Optimizing Prompts for Text-to-Image Generation

Yaru HaoZewen ChiLi DongFuru Wei

article2023NeurIPS293 citations

Proposes PROMPTIST, a framework combining supervised fine-tuning and reinforcement learning to automatically rewrite plain text inputs into optimized prompts that boost image aesthetics in text-to-image models while preserving user intent.

Listen

Text-to-image artificial intelligence systems often require complex, highly customized text prompts to produce high-quality visual outputs. Because ordinary users typically supply natural, concise descriptions, their inputs frequently lead to suboptimal images. Relying on manual prompt engineering—such as manually appending artist names and stylistic tags—is labor-intensive, difficult for non-technical users, and largely non-transferable across different model versions.

The article demonstrates and evaluates an automated framework called PROMPTIST, designed to bridge this gap by translating simple user inputs into optimized, model-preferred prompts that enhance image aesthetics while strictly preserving the user's original intent.

The researchers developed a two-stage training approach using a 117-million-parameter language model. First, the model underwent supervised fine-tuning on 360,000 prompt pairs derived from human-engineered prompt galleries and automated rephrasings. Second, the system applied reinforcement learning via Proximal Policy Optimization to explore improved prompt variations across hundreds of thousands of standard and out-of-domain text inputs. The reward signal evaluated generated images from Stable Diffusion on two core dimensions: aesthetic appeal and relevance to the original user input, balanced by a penalty to prevent erratic modifications.

Evaluation results show that the automated framework substantially outperformed both original user inputs and manual prompt engineering. In automated scoring, reinforcement learning achieved significant gains over supervised fine-tuning alone, improving rewards by 24% to 31% on in-domain inputs and by 71% on out-of-domain captions. On Stable Diffusion, the method maintained strong relevance to the user's concept while boosting the aesthetic rating from 5.47 (raw input) and 5.87 (human-engineered prompt) up to 6.26. In human evaluation studies, annotators preferred the images generated from automated prompts over raw user inputs 72% of the time, and preferred them over laboriously human-engineered prompts 49% of the time, compared to a 21% preference for human prompts and 30% ties.

These findings indicate that language models can serve as an efficient, automated translation layer between end users and generative visual tools. This reduces the time, cost, and expertise required to operate text-to-image systems, making advanced generative models accessible to non-experts without degrading visual quality.

Organizations deploying generative image tools should consider integrating automated prompt translation layers rather than relying on manual user training or static prompt cheat sheets. For next steps, the framework should be tested on direct human-in-the-loop feedback mechanisms and expanded into other modalities, such as text-to-video generation.

Decision-makers should note certain limitations: the training data contains inherent stylistic biases toward digital art rather than photorealism and features a high proportion of portraits. While confidence in the framework's effectiveness across evaluated benchmarks is high, results should be carefully piloted in settings requiring realistic photography or strict adherence to specialized corporate brand standards.

arXiv: 2212.09611
Cover for Optimizing Prompts for Text-to-Image Generation

Abstract

Well-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation, a general framework that automatically adapts original user input to model-preferred prompts. Specifically, we first perform supervised fine-tuning with a pretrained language model on a small collection of manually engineered prompts. Then we use reinforcement learning to explore better prompts. We define a reward function that encourages the policy to generate more aesthetically pleasing images while preserving the original user intentions. Experimental results on Stable Diffusion show that our method outperforms manual prompt engineering in terms of both automatic metrics and human preference ratings. Moreover, reinforcement learning further boosts performance, especially on out-of-domain prompts. The pretrained checkpoints are available at https://aka.ms/promptist. The demo can be found at https://aka.ms/promptist-demo.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Supervised fine-tuning
  • 2.2 Reward definition
  • 2.3 Reinforcement learning
  • 3 Experiments
  • 3.1 Data collection
  • 3.2 Settings
  • 3.3 Results
  • 3.4 Human evaluation
  • 3.5 Ablation of source prompt augmentation
  • 4 Related work
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgments
  • References
  • Appendix
  • A Hyperparameter settings
  • B Computational budget
  • C Results on Stable Diffusion v1.5.
  • D Comparisons with heuristic baseline
  • E Results on different categories and lengths of prompts

Knowls

  1. Knowl 1 — Prompt Adaptation Framework for Text-to-Image Generation (PROMPTIST)

    model/method

    PROMPTIST is an automatic prompt adaptation framework that converts plain, user-provided text inputs into model-preferred prompts to improve the visual aesthetic quality of text-to-image diffusion models (such as Stable Diffusion) while maintaining original semantic intent. The framework functions as an intermediate language model interface optimized in two successive phases:

    1. Supervised Fine-Tuning (SFT): A pretrained autoregressive language model (GPT-2, 117M parameters) is fine-tuned on paired prompt data mapping simplified or rephrased user inputs to human-engineered prompts containing effective stylistic modifiers.
    2. Reinforcement Learning (RL): The fine-tuned policy model is further trained using Proximal Policy Optimization (PPO) by generating candidate prompts and receiving image-level feedback from the frozen text-to-image generator. The reward function couples an automatic aesthetic quality score and a CLIP-based semantic relevance score, along with a Kullback-Leibler (KL) divergence penalty against the supervised baseline to prevent reward over-optimization.
  2. Knowl 2 — Parallel Prompt Corpus Construction and Supervised Fine-Tuning

    model/method

    Supervised fine-tuning of the prompt adaptation policy requires parallel pairs (x,y)(x, y), where xx is an unembellished input prompt and yy is a performant target prompt. Using 90,000 human-engineered target prompts crawled from the online gallery Lexica, four source prompt variations are synthesized for each target, yielding 360,000 paired training examples:

    1. Main Content Extraction (MC): Stylistic modifiers (e.g., artist names, resolution buzzwords, rendering engines) are removed, isolating the central semantic subject.
    2. Main Content with Random Modifiers (MCM): Modifiers from the target prompt are randomly shuffled or partially deleted.
    3. Rephrasing of Main Content (RMC): Extracted main contents are rephrased into natural, conversational language using an LLM (text-davinci-002) with the instruction prompt "[Input] Rephrase:".
    4. Rephrasing of Target Prompt (RTP): The complete engineered target prompt is rephrased into conversational sentences via text-davinci-002.

    The policy model parameters θ\theta are trained using teacher forcing to minimize the sequence cross-entropy loss:

    LSFT=−E(x,y)∼D[∑t=1∣y∣log⁡pθ(yt∣x,y<t)]\mathcal{L}_{\text{SFT}} = -\mathbb{E}_{(x, y) \sim \mathcal{D}} \left[ \sum_{t=1}^{|y|} \log p_\theta(y_t \mid x, y_{<t}) \right]

    where D\mathcal{D} is the augmented parallel dataset.

  3. Knowl 3 — Reinforcement Learning Total Reward Formulation

    equation

    In the reinforcement learning stage, the policy πθ\pi_\theta generates an adapted prompt yy conditioned on an input prompt xx. The objective maximizes the expected cumulative reward across prompt distributions. The total scalar reward R(x,y)R(x, y) combines an aesthetic score faes(x,y)f_{\text{aes}}(x, y), a relevance score frel(x,y)f_{\text{rel}}(x, y), and a Kullback-Leibler (KL) divergence penalty against the supervised fine-tuned model πSFT\pi_{\text{SFT}}:

    R(x,y)=faes(x,y)+frel(x,y)−ηlog⁡πθ(y∣x)πSFT(y∣x)R(x, y) = f_{\text{aes}}(x, y) + f_{\text{rel}}(x, y) - \eta \log \frac{\pi_\theta(y \mid x)}{\pi_{\text{SFT}}(y \mid x)}

    where:

    • x=(x1,…,xn)∈Vnx = (x_1, \dots, x_n) \in \mathcal{V}^n is the input prompt token sequence over vocabulary V\mathcal{V}.
    • y=(y1,…,ym)∈Vmy = (y_1, \dots, y_m) \in \mathcal{V}^m is the generated prompt token sequence.
    • faes(x,y)∈Rf_{\text{aes}}(x, y) \in \mathbb{R} evaluates the visual aesthetic gain of images generated with yy over images generated with xx.
    • frel(x,y)≤0f_{\text{rel}}(x, y) \le 0 penalizes deviations between images generated with yy and the original intent xx.
    • πSFT(y∣x)\pi_{\text{SFT}}(y \mid x) is the conditional token probability from the frozen supervised fine-tuned model.
    • η∈R+\eta \in \mathbb{R}^+ is the KL penalty coefficient (set to η=0.2\eta = 0.2) used to restrict policy drift.
  4. Knowl 4 — Aesthetic and Semantic Relevance Reward Metrics

    equation

    Let G(p)G(p) represent the text-to-image synthesis model generating an image conditioned on prompt pp, let gCLIP(x,i)g_{\text{CLIP}}(x, i) denote the cosine similarity between the CLIP text embedding of user prompt xx and the CLIP image embedding of image ii, and let gaes(i)g_{\text{aes}}(i) denote the aesthetic quality predictor (a linear regression head trained on frozen CLIP image embeddings using human ratings from the AVA dataset).

    The semantic relevance score frel(x,y)f_{\text{rel}}(x, y) assesses whether images conditioned on adapted prompt yy retain the meaning of original prompt xx:

    frel(x,y)=Eiy∼G(y)[frel(x,iy)]f_{\text{rel}}(x, y) = \mathbb{E}_{i_y \sim G(y)} [f_{\text{rel}}(x, i_y)]

    frel(x,iy)=min⁡(20⋅gCLIP(x,iy)−5.6, 0)f_{\text{rel}}(x, i_y) = \min(20 \cdot g_{\text{CLIP}}(x, i_y) - 5.6,\, 0)

    When the CLIP similarity gCLIP(x,iy)≥0.28g_{\text{CLIP}}(x, i_y) \ge 0.28, frel(x,iy)=0f_{\text{rel}}(x, i_y) = 0, giving no penalty. When similarity falls below 0.280.28, a linearly decreasing negative penalty is applied.

    The aesthetic reward faes(x,y)f_{\text{aes}}(x, y) measures relative aesthetic improvement over the raw user prompt:

    faes(x,y)=Eix∼G(x),iy∼G(y)[gaes(iy)−gaes(ix)]f_{\text{aes}}(x, y) = \mathbb{E}_{i_x \sim G(x), i_y \sim G(y)} [g_{\text{aes}}(i_y) - g_{\text{aes}}(i_x)]

    Because gCLIPg_{\text{CLIP}} and gaesg_{\text{aes}} both use the CLIP vision encoder, image representations are computed in a single shared forward pass.

  5. Knowl 5 — RL Policy Optimization and Exploration Dynamics

    experimental setup

    The prompt adaptation policy πθ\pi_\theta and value network VϕV_\phi are initialized from the SFT GPT-2 model (117M parameters) with separate parameter sets. The model is optimized using Proximal Policy Optimization (PPO) with clipped surrogate objectives for 12,000 episodes over unlabelled source prompts:

    1. In-domain real user prompts from DiffusionDB (600,000 examples).
    2. Out-of-domain image descriptions from MS-COCO (600,000 examples).
    3. Short class labels from ImageNet-21k (40,000 examples).

    Exploration and stability techniques include:

    • Diverse Beam Search: Prompt generation employs Diverse Beam Search (beam size 8, diversity penalty 1.0), and one candidate completion is sampled randomly from the diverse beam set to update policy parameters.
    • Dynamic Generation Length: To prevent the model from over-generating lengthy modifiers for short inputs, the maximum decoding length per step is randomly sampled from U{15,75}\mathcal{U}\{15, 75\}.
    • Reward Variance Reduction: For each candidate prompt, Stable Diffusion v1.4 generates 3 images using the DPM solver (20 denoising steps), and scores are averaged.
    • Hyperparameters: Batch size 256, learning rate 5×10−55 \times 10^{-5} (Adam optimizer), value loss coefficient 2.3, KL coefficient 0.2, and 4 PPO epochs per batch with one minibatch each.
  6. Knowl 6 — Reward Gains of Reinforcement Learning Over Supervised Fine-Tuning

    data/table

    The performance of adapted prompts was evaluated on held-out test splits (256 prompts per category) generated via Stable Diffusion v1.4. The table displays absolute reward improvements over raw user inputs for Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL).

    Model In-Domain (Lexica) In-Domain Out-of-Domain
    MC MCM RMC RTP (DiffusionDB) (COCO)
    SFT 0.36 0.16 0.44 0.11 0.29 0.28
    RL 0.47 0.17 0.63 0.25 0.36 0.48
    Gain +31% +6% +43% +127% +24% +71%

    The categories correspond to:

    • MC: Main content.
    • MCM: Main content with random modifiers.
    • RMC: Rephrasing of main content.
    • RTP: Rephrasing of target prompt.
    • DiffusionDB: In-domain user prompts.
    • COCO: Out-of-domain captions.

    While supervised fine-tuning provides initial gains, RL exploration yields large relative improvements, particularly on conversational rephrasings (+43% on RMC, +127% on RTP) and out-of-domain COCO prompts (+71%).

  7. Knowl 7 — Aesthetic Score and CLIP Relevance on DiffusionDB

    data/table

    To evaluate trade-offs between aesthetic enhancement and semantic fidelity, aesthetic predictor scores and CLIP text-image cosine similarities were measured on DiffusionDB test queries using Stable Diffusion v1.4.

    Prompt Source Aesthetic Score Relevance Score
    User Input 5.47 0.28
    Human Engineered Prompt 5.87 0.26
    Supervised Fine-Tuning (SFT) 6.15 0.25
    PROMPTIST (Ours) 6.26 0.26

    A CLIP relevance score around 0.26 indicates sufficient semantic fidelity. PROMPTIST achieves an aesthetic score of 6.26, exceeding both human-engineered prompts (5.87) and the SFT baseline (6.15), while preserving relevance at 0.26 (matching human-engineered prompts and improving upon SFT's 0.25).

  8. Knowl 8 — Human Preference Evaluation of PROMPTIST Prompts

    data/table

    Human evaluation was conducted with three independent annotators ranking image pairs generated by Stable Diffusion v1.4 from raw user inputs, human-engineered prompts, and PROMPTIST-adapted prompts (2 images per prompt).

    Comparison PROMPTIST Preferred Equal Preference Baseline Preferred
    In-Domain: PROMPTIST vs. User Input 72% 13% 15%
    Out-of-Domain: PROMPTIST vs. User Input 67% 12% 21%
    In-Domain: PROMPTIST vs. Manually Engineered 49% 21% 30%

    Annotators favored images generated from PROMPTIST prompts over raw user inputs on both in-domain queries (72% vs. 15%) and out-of-domain descriptions (67% vs. 21%). Compared against labor-intensive manual prompt engineering, PROMPTIST prompts achieved a 49% preference rate versus 30% for human prompts, with 21% rated equal.

  9. Knowl 9 — Model Transferability and Heuristic Modifier Baseline Comparison

    empirical result

    Evaluation of PROMPTIST across generators and against rule-based modifier baselines shows:

    1. Generalization to Stable Diffusion v1.5: PROMPTIST consistently outperforms user inputs, human engineering, and SFT without retraining on v1.5. On test rewards: Lexica (+0.05 vs. -0.05 human, -0.31 user), DiffusionDB (+0.06 vs. -0.18 human, -0.32 user), and COCO (+0.11 vs. -0.10 SFT, -0.37 user).
    2. Failure of Static Heuristic Tags: Concatenating popular static tag combinations (e.g., "artstation, highly detailed, elegant" or "digital painting, intricate, fantasy") from the top 15 most frequent Lexica keywords produces inconsistent rewards across domains (e.g., the fantasy tag achieves +0.06 on Lexica and -0.20 on COCO, but drops to -0.16 on DiffusionDB). PROMPTIST consistently generates higher, positive rewards across all datasets (+0.14 on Lexica, +0.06 on DiffusionDB, +0.10 on COCO) by dynamically selecting modifiers tailored to each prompt.
  10. Knowl 10 — Source Prompt Augmentation Ablation in Supervised Fine-Tuning

    empirical result

    Ablation of source prompt augmentation during supervised fine-tuning indicates:

    • Fine-tuning GPT-2 strictly on extracted main content without augmentation (no random modifier perturbations and no LLM rephrasings) leads to poor generalization on conversational and out-of-domain prompt distributions.
    • Including the four constructed source prompt types (MC, MCM, RMC, RTP) yields consistent reward gains across both in-domain held-out test splits and out-of-domain COCO captions.
    • Manually adding random modifiers to user input (the MCM baseline) produces minor improvements compared to fine-tuning, demonstrating that effective prompt adaptation requires context-specific modifier selection rather than arbitrary keyword addition.
  11. Knowl 11 — Dataset Biases in Crawled Human-Engineered Prompts

    limitation

    The supervised training data constructed from Lexica exhibits two primary distributional biases:

    1. Artistic Style Bias: Crawled prompts frequently feature specific digital artist names and stylistic rendering tags (e.g., Greg Rutkowski, ArtStation), biasing generated outputs toward digital artwork and fantasy illustrations rather than realistic photography.
    2. Subject Imbalance: Portrait and character prompts represent a disproportionately large share of the dataset relative to other categories such as natural landscapes, architecture, or everyday objects.

    While reinforcement learning exploration across open-domain caption datasets mitigates these tendencies, balancing subject and style distributions during initial demonstration collection remains necessary for domain-general prompt adaptation.

Coverage note — None omitted; the knowls cover the two-stage prompt adaptation framework, data construction methodology, reward formulations, RL exploration and training setup, empirical reward improvements, aesthetic/relevance evaluations, human preference studies, heuristic baseline comparisons, ablations, and dataset limitations.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020.
  2. 2.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO captions: Data collection and evaluation server. 2015.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek B Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Oliveira Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. PaLM: Scaling language modeling with pathways. ArXiv, abs/2204.02311, 2022.
  4. 4.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  5. 5.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.581. URL https://aclanthology.org/2022.acl-long.581.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  7. 7.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems, 34:19822–19835, 2021.
  8. 8.Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. Glam: Efficient scaling of language models with mixture-of-experts, 2021.
  9. 9.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.295. URL https://aclanthology.org/2021.acl-long.295.
  10. 10.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022.
  11. 11.Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. Learning from dialogue after deployment: Feed yourself, chatbot! In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3667–3684, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1358. URL https://aclanthology.org/P19-1358.
  12. 12.Adi Haviv, Jonathan Berant, and Amir Globerson. BERTese: Learning to speak to BERT. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3618–3623, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.316. URL https://aclanthology.org/2021.eacl-main.316.
  13. 13.Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. Reward learning from human preferences and demonstrations in atari. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper/2018/file/8cbe9ce23f42628c98f80fa0fac8b19a-Paper.pdf.
  14. 14.Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456, 2019.
  15. 15.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. How Can We Know What Language Models Know? Transactions of the Association for Computational Linguistics, 8:423–438, 07 2020. ISSN 2307-387X. doi: 10.1162/tacl_a_00324. URL https://doi.org/10.1162/tacl_a_00324.
  16. 16.Jan-Christoph Klie, Richard Eckart de Castilho, and Iryna Gurevych. From Zero to Hero: Human-In-The-Loop Entity Linking in Low Resource Domains. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6982–6993, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.624. URL https://aclanthology.org/2020.acl-main.624.
  17. 17.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.353. URL https://aclanthology.org/2021.acl-long.353.
  18. 18.Vivian Liu and Lydia B. Chilton. Design guidelines for prompt engineering text-to-image generative models, 2021.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  20. 20.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  21. 21.Naila Murray, Luca Marchesotti, and Florent Perronnin. Ava: A large-scale database for aesthetic visual analysis. In 2012 IEEE conference on computer vision and pattern recognition, pages 2408–2415. IEEE, 2012.
  22. 22.Jonas Oppenlaender. A taxonomy of prompt modifiers for text-to-image generation, 2022.
  23. 23.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  24. 24.Guy Parsons. The dall·e 2 prompt book, 2022.
  25. 25.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1250. URL https://aclanthology.org/D19-1250.
  26. 26.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/radford21a.html.
  28. 28.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. arXiv, abs/2102.12092, 2021a.
  29. 29.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021b.
  30. 30.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022.
  31. 31.Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis. In International conference on machine learning, pages 1060–1069. PMLR, 2016a.
  32. 32.Scott E Reed, Zeynep Akata, Santosh Mohan, Samuel Tenka, Bernt Schiele, and Honglak Lee. Learning what and where to draw. Advances in neural information processing systems, 29, 2016b.
  33. 33.Laria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm, 2021.
  34. 34.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022.
  35. 35.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding, 2022.
  36. 36.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  37. 37.Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. Constrained language models yield few-shot semantic parsers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7699–7715, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.608. URL https://aclanthology.org/2021.emnlp-main.608.
  38. 38.Kurt Shuster, Jack Urbanek, Emily Dinan, Arthur Szlam, and Jason Weston. Deploying lifelong open-domain dialogue learning. arXiv preprint arXiv:2008.08076, 2020.
  39. 39.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model, 2022.
  40. 40.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  41. 41.Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu. Df-gan: A simple and effective baseline for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16515–16525, 2022.
  42. 42.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212, 2021.
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008. Curran Associates, Inc., 2017. URL http://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf.
  44. 44.Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun, Stefan Lee, David J. Crandall, and Dhruv Batra. Diverse beam search: Decoding diverse solutions from neural sequence models. ArXiv, abs/1610.02424, 2016.
  45. 45.Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. DiffusionDB: A large-scale prompt gallery dataset for text-to-image generative models. arXiv:2210.14896 [cs], 2022. URL https://arxiv.org/abs/2210.14896.
  46. 46.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=_VjQlMeSB_J.
  47. 47.Jing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora, Y-Lan Boureau, and Jason Weston. Learning new skills after deployment: Improving open-domain internet-driven dialogue with human feedback. arXiv preprint arXiv:2208.03270, 2022.
  48. 48.Ziyu Yao, Yu Su, Huan Sun, and Wen-tau Yih. Model-based interactive semantic parsing: A unified framework and a text-to-SQL case study. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5447–5458, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1547. URL https://aclanthology.org/D19-1547.
  49. 49.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022a.
  50. 50.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large language models are human-level prompt engineers, 2022b.
  51. 51.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Hao, Y., et al. “Optimizing Prompts for Text-to-Image Generation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 66923–39, https://proceedings.neurips.cc/paper_files/paper/2023/file/d346d91999074dd8d6073d4c3b13733b-Paper-Conference.pdf.
APA
Hao, Y., Chi, Z., Dong, L., & Wei, F. (2023). Optimizing Prompts for Text-to-Image Generation. Advances in Neural Information Processing Systems, 36, 66923–66939. https://proceedings.neurips.cc/paper_files/paper/2023/file/d346d91999074dd8d6073d4c3b13733b-Paper-Conference.pdf
Chicago
Hao, Y., Z. Chi, L. Dong, and F. Wei. 2023. “Optimizing Prompts for Text-to-Image Generation”. Advances in Neural Information Processing Systems 36: 66923–39. https://proceedings.neurips.cc/paper_files/paper/2023/file/d346d91999074dd8d6073d4c3b13733b-Paper-Conference.pdf.
Harvard
Hao, Y. et al. (2023) “Optimizing Prompts for Text-to-Image Generation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 66923–66939. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/d346d91999074dd8d6073d4c3b13733b-Paper-Conference.pdf.
Vancouver
1. Hao Y, Chi Z, Dong L, Wei F (2023) Optimizing Prompts for Text-to-Image Generation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 66923–66939

BibTeX

@inproceedings{hao2023optimizing,
  title = {Optimizing Prompts for Text-to-Image Generation},
  author = {Hao, Yaru and Chi, Zewen and Dong, Li and Wei, Furu},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {66923-66939},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/d346d91999074dd8d6073d4c3b13733b-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission