Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Yuval KirstainAdam PolyakUriel SingerShahbuland MatianaJoe PennaOmer Levy

article2023NeurIPS864 citations

Introduces an open dataset of real user preferences for text-to-image synthesis alongside PickScore, a scoring function that predicts human judgment more accurately than existing metrics to improve automated evaluation and model ranking.

Listen

Aligning text-to-image artificial intelligence models with actual user preferences is critical for creating high-quality generative tools. However, progress has been constrained because large-scale preference datasets remain proprietary and locked within private corporations. The article addresses this gap by developing an open dataset of real user preferences, creating an automated scoring function trained on this data, and establishing a more accurate benchmark for evaluating and improving text-to-image models.

The authors built an interactive web application that allowed real, intrinsically motivated users to generate images from custom prompts and select their preferred output. Through this platform, they gathered over 500,000 preference pairs spanning 35,000 distinct prompts. Leveraging this open resource, named Pick-a-Pic, the authors fine-tuned an image-text scoring model called PickScore to estimate how satisfied a user is with a generated image relative to their prompt.

Key findings demonstrate the superiority of authentic user data over conventional metrics. First, PickScore achieved a 70.5% accuracy rate in predicting user choices on held-out test data, outperforming both human expert annotators (68.0%) and existing baselines like standard CLIP-H (60.8%) and aesthetic predictors (56.8%). Second, traditional automated benchmarks were shown to be deeply flawed; for example, the widely used Fréchet Inception Distance metric exhibited a strong negative correlation (-0.900) with human rankings on standard image captions, whereas PickScore showed a strong positive correlation (0.917). Third, when evaluating models against actual user preferences, PickScore's rankings achieved a 0.790 correlation with ground truth, significantly outperforming alternative scoring metrics. Finally, using PickScore to automatically select the best image out of 100 generated variations yielded human preference win rates exceeding 71% against unranked model outputs and competing scoring functions.

These results carry significant strategic implications for machine learning development. Traditional evaluation pipelines rely on photographic caption datasets like MS-COCO, which do not reflect the creative, fictional prompts real users actually write. Furthermore, optimizing models against outdated realism metrics can inadvertently degrade the vivid, aesthetically pleasing qualities users prefer. Incorporating a robust automated judge like PickScore provides an efficient, low-cost way to boost generation quality through post-processing rank selection or direct model alignment (such as reinforcement learning from human feedback) without requiring continuous human labeling.

Organizations developing or deploying text-to-image systems should immediately replace or supplement standard benchmarks with Pick-a-Pic prompts and adopt PickScore for automated performance assessments. When deploying generative applications, engineering teams should evaluate using PickScore as a reranking filter over candidate images to improve output quality. Decision-makers should note that while the dataset underwent strict filtering, some residual not-safe-for-work content and user bias may remain. Nonetheless, the high statistical consistency across extensive trials provides strong confidence in PickScore as a primary evaluation and selection tool.

Cover for Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Abstract

The ability to collect a large dataset of human preferences from text-to-image users is usually limited to companies, making such datasets inaccessible to the public. To address this issue, we create a web app that enables text-to-image users to generate images and specify their preferences. Using this web app we build Pick-a-Pic, a large, open dataset of text-to-image prompts and real users' preferences over generated images. We leverage this dataset to train a CLIP-based scoring function, PickScore, which exhibits superhuman performance on the task of predicting human preferences. Then, we test PickScore's ability to perform model evaluation and observe that it correlates better with human rankings than other automatic evaluation metrics. Therefore, we recommend using PickScore for evaluating future text-to-image generation models, and using Pick-a-Pic prompts as a more relevant dataset than MS-COCO. Finally, we demonstrate how PickScore can enhance existing text-to-image models via ranking.

Table of Contents

  • 1 Introduction
  • 2 Pick-a-Pic Dataset
  • 3 PickScore
  • 4 Preference Prediction
  • 5 Model Evaluation
  • 6 Text-to-Image Ranking
  • 7 Related Work
  • 8 Limitations and Broader Impact
  • 9 Conclusions
  • References

Knowls

  1. Knowl 1 — PickScore Architecture and Preference Objective

    model/method

    PickScore is a scoring function designed to predict human preference for a generated image yy conditioned on a text prompt xx. It adopts the dual-encoder architecture of CLIP, computing a scalar compatibility score s(x,y)s(x, y) as the inner product of normalized transformer text and image embeddings scaled by a learned temperature parameter TT:

    s(x,y)=Etxt(x)⋅Eimg(y)⋅Ts(x, y) = E_{\text{txt}}(x) \cdot E_{\text{img}}(y) \cdot T

    To train the scoring function on pairwise user preferences, the model minimizes the Kullback-Leibler (KL) divergence between a target human preference distribution pp and the softmax-normalized scores p^\hat{p} over a pair of candidate images (y1,y2)(y_1, y_2) generated for prompt xx:

    p^i=exp⁡s(x,yi)∑j=12exp⁡s(x,yj)\hat{p}_i = \frac{\exp s(x, y_i)}{\sum_{j=1}^2 \exp s(x, y_j)}

    Lpref=∑i=12pi(log⁡pi−log⁡p^i)\mathcal{L}_{\text{pref}} = \sum_{i=1}^2 p_i (\log p_i - \log \hat{p}_i)

    where the ground-truth distribution is p=[1,0]p = [1, 0] if y1y_1 is preferred, p=[0,1]p = [0, 1] if y2y_2 is preferred, or p=[0.5,0.5]p = [0.5, 0.5] in the case of a tie. When aggregating the loss across a batch, each sample is weighted inversely proportional to the frequency of its prompt in the entire dataset to prevent overfitting to heavily repeated prompts.

    PickScore is initialized by fine-tuning OpenCLIP ViT-H (CLIP-H) for 4,000 steps with a batch size of 128, a learning rate of 3×10−63 \times 10^{-6}, and a 500-step warmup followed by linear decay.

  2. Knowl 2 — Pick-a-Pic Dataset Construction and Split Strategy

    experimental setup

    The Pick-a-Pic dataset is an open collection of text-to-image prompts and authentic user pairwise preferences collected via a public web application. In the application, real users write prompts and iteratively compare pairs of generated images, selecting their preferred image or declaring a tie; the non-preferred image is replaced with a newly generated image in subsequent rounds. The images are generated using multiple generative backbones—including Stable Diffusion 2.1, Dreamlike Photoreal 2.0 (a fine-tuned Stable Diffusion 1.5 variant), and Stable Diffusion XL variants—sampled across various classifier-free guidance scale values.

    Quality control and safety measures include requiring Gmail or Discord user authentication, filtering NSFW keywords and banning abusive accounts, and capping single-user activity.

    To ensure unbiased evaluation, dataset splits are constructed without prompt overlap: 1,000 distinct prompts from 1,000 unique users are sampled and partitioned equally into validation and test sets (500 examples each, with exactly one preference pair per prompt). The remaining examples form the training split (comprising 583,747 pairwise comparisons across 37,523 distinct prompts from 4,375 users in the primary experimental release).

  3. Knowl 3 — Tie-Aware Evaluation Metric and Threshold Selection for Pairwise Preference

    model/method

    To evaluate binary scoring models on human preference prediction tasks that include ties, an adapted scoring metric assigns:

    • 1.0 point if the model's predicted outcome matches the user judgment exactly (image 1, image 2, or tie);
    • 0.5 points if either the user label or the model prediction is a tie (while the other indicates a strict preference);
    • 0.0 points if the prediction and user label are opposing strict preferences, or if one predicted a tie while the other was completely discordant.

    To produce tie predictions from a model yielding normalized softmax probabilities p^1\hat{p}_1 and p^2\hat{p}_2, a decision threshold t∈[0,1]t \in [0, 1] is established such that the model predicts a tie whenever:

    ∣p^1−p^2∣<t|\hat{p}_1 - \hat{p}_2| < t

    The threshold tt is individually tuned on the validation set for each model to maximize its accuracy under this metric before evaluating on the test set.

  4. Knowl 4 — User Preference Prediction Benchmark Performance

    data/table

    When evaluated on the held-out Pick-a-Pic test set using the tie-aware preference accuracy metric, PickScore outperforms expert human annotators as well as prior and concurrent automated scoring models.

    Model Accuracy (%)
    Random Baseline 56.8
    LAION Aesthetics Predictor 56.8
    Zero-shot CLIP-H 60.8
    ImageReward 61.1
    Human Preference Score (HPS) 66.7
    Human Expert Annotators 68.0
    PickScore (Mean ±\pm Variance) 70.5 ±\pm 0.142

    The gap between human expert annotators (68.0%) and PickScore (70.5%) reflects that external annotators lack access to the original user's latent creative intent and unstated context, whereas PickScore successfully generalizes real user preference patterns across prompts.

  5. Knowl 5 — FID Negative Correlation with Human Preferences Under Classifier-Free Guidance

    empirical result

    Fréchet Inception Distance (FID) on MS-COCO validation captions fails to reflect human preferences when evaluating image generation across different classifier-free guidance (CFG) scales. When 9 model configurations (Stable Diffusion 1.5, Stable Diffusion 2.1, and Dreamlike Photoreal 2.0 at CFG scales 3, 6, and 9) are evaluated on 100 MS-COCO validation captions:

    • Model win rates ranked by FID exhibit a strong negative Spearman rank correlation with human expert preferences: ρ=−0.900\rho = -0.900.
    • Model win rates ranked by PickScore exhibit a strong positive Spearman rank correlation with human expert preferences: ρ=0.917\rho = 0.917.

    This discrepancy arises because higher classifier-free guidance scales produce more vivid, detailed images that human users strongly prefer, but shift the feature representation away from the distribution of amateur MS-COCO photographs, resulting in worse (higher) FID scores despite superior perceived visual quality.

  6. Knowl 6 — Alignment with Real-User Model Elo Ratings

    empirical result

    When evaluating text-to-image models using Elo ratings calculated over 14,000 real-user preference judgments on the Pick-a-Pic test set across 45 model configurations (comprising 4 distinct backbone architectures evaluated across various guidance scales), PickScore demonstrates substantially higher correlation with ground-truth user Elo rankings than competing automatic evaluation metrics:

    • Zero-shot CLIP-H: Spearman ρ=0.313±0.075\rho = 0.313 \pm 0.075
    • ImageReward: Spearman ρ=0.492±0.086\rho = 0.492 \pm 0.086
    • Human Preference Score (HPS): Spearman ρ=0.670±0.071\rho = 0.670 \pm 0.071
    • PickScore: Spearman ρ=0.790±0.054\rho = 0.790 \pm 0.054

    Correlations and standard deviations were computed over 50 randomized ordering iterations of the iterative Elo calculation.

  7. Knowl 7 — Text-to-Image Generation Improvement via PickScore Best-of-N Ranking

    data/table

    PickScore can improve the output quality of vanilla text-to-image models via post-hoc candidate ranking. For each test prompt, 100 candidate images are generated using Dreamlike Photoreal 2.0 (CFG 7.5) across 5 random initial noise seeds and 20 prompt augmentation templates (e.g., prepending/appending quality modifiers such as 'breathtaking [prompt]. award-winning, professional, highly detailed'). Human raters evaluated the top candidate selected by PickScore against selections made by alternative scoring functions.

    Comparison PickScore Win Rate (%)
    PickScore vs. Random Seed + Null Template 71.4
    PickScore vs. Random Seed + Random Template 82.0
    PickScore vs. Aesthetics Predictor Selection 85.1
    PickScore vs. Zero-shot CLIP-H Selection 71.3

    PickScore achieves balanced selection: in 68.5% of prompts, PickScore selects an image with a higher aesthetic score than CLIP-H, and in 90.5% of prompts, it selects an image with higher text-alignment than the pure aesthetics predictor.

  8. Knowl 8 — Degradation of Preference Modeling from In-Batch Negatives

    empirical result

    Modifying the preference loss objective to include in-batch negatives across all candidate images in the batch—formulating the normalized score for candidate ii of the kk-th sample as:

    p^ik=exp⁡s(x,yi)∑m∑j=12exp⁡s(x,yjm)\hat{p}_i^k = \frac{\exp s(x, y_i)}{\sum_m \sum_{j=1}^2 \exp s(x, y_j^m)}

    where yjmy_j^m denotes image jj belonging to example mm in the batch—degrades human preference prediction performance. Training CLIP-H with this in-batch negative objective yields a Pick-a-Pic test set accuracy of 65.2%, compared to 70.5% achieved by training strictly on the localized pairwise preference objective.

  9. Knowl 9 — Limitations and Potential Biases of Pick-a-Pic Data

    limitation

    The Pick-a-Pic dataset and models trained on it possess several inherent limitations:

    1. Residual NSFW content: Despite automated filtering of sensitive keywords and user bans, some NSFW text and imagery remain in the collected dataset.
    2. Rater inattention and noise: Because interactions come from open web users rather than vetted annotators under supervised conditions, some comparisons may be clicked carelessly or at high speeds.
    3. User demographic biases: Preferences reflect the specific population of users participating in online text-to-image communities, which may incorporate cultural, demographic, or aesthetic biases that differ from the general population.

Coverage note — None was omitted; all key contributions including dataset construction, PickScore architecture, objective formulation, preference prediction benchmarks, model evaluation alignment against FID/Elo, Best-of-N ranking experiments, negative results on in-batch negatives, and stated limitations are represented.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, T. J. Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv, abs/2204.05862, 2022.
  2. 2.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. ArXiv, abs/1504.00325, 2015.
  3. 3.Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. ArXiv, abs/1706.03741, 2017.
  4. 4.Arpad E. Elo. The rating of chessplayers, past and present. 1978.
  5. 5.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS, 2017.
  6. 6.Jonathan Ho. Classifier-free diffusion guidance. ArXiv, abs/2207.12598, 2022.
  7. 7.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip, July 2021. If you use this software, please cite it as below.
  8. 8.Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, P. Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu. Aligning text-to-image models using human feedback. ArXiv, abs/2302.12192, 2023.
  9. 9.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  10. 10.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  11. 11.John David Pressman, Katherine Crowson, and Simulacra Captions Contributors. Simulacra aesthetic captions. Technical Report Version 1.0, Stability AI, 2022. url https://github.com/JD-P/simulacra-aesthetic-captions .
  12. 12.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021.
  13. 13.Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, 2021.
  14. 14.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5b: An open large-scale dataset for training next generation image-text models. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  15. 15.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, 2015.
  16. 16.Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. ArXiv, abs/2210.14896, 2022.
  17. 17.Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li. Better aligning text-to-image models with human preference. ArXiv, abs/2303.14420, 2023.
  18. 18.Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagereward: Learning and evaluating human preferences for text-to-image generation. ArXiv, abs/2304.05977, 2023.

Citation

MLA
Kirstain, Y., et al. “Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation”. arXiv, 2023, http://arxiv.org/abs/2305.01569v2.
APA
Kirstain, Y., Polyak, A., Singer, U., Matiana, S., Penna, J., & Levy, O. (2023). Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. arXiv. http://arxiv.org/abs/2305.01569v2
Chicago
Kirstain, Y., A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy. 2023. “Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation”. arXiv. http://arxiv.org/abs/2305.01569v2.
Harvard
Kirstain, Y. et al. (2023) “Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.01569v2.
Vancouver
1. Kirstain Y, Polyak A, Singer U, Matiana S, Penna J, Levy O (2023) Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. arXiv

BibTeX

@article{kirstain2023pick,
  title = {Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation},
  author = {Kirstain, Yuval and Polyak, Adam and Singer, Uriel and Matiana, Shahbuland and Penna, Joe and Levy, Omer},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.01569v2},
  eprint = {2305.01569}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors