Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
Yuval KirstainAdam PolyakUriel SingerShahbuland MatianaJoe PennaOmer Levy
Introduces an open dataset of real user preferences for text-to-image synthesis alongside PickScore, a scoring function that predicts human judgment more accurately than existing metrics to improve automated evaluation and model ranking.
Aligning text-to-image artificial intelligence models with actual user preferences is critical for creating high-quality generative tools. However, progress has been constrained because large-scale preference datasets remain proprietary and locked within private corporations. The article addresses this gap by developing an open dataset of real user preferences, creating an automated scoring function trained on this data, and establishing a more accurate benchmark for evaluating and improving text-to-image models.
The authors built an interactive web application that allowed real, intrinsically motivated users to generate images from custom prompts and select their preferred output. Through this platform, they gathered over 500,000 preference pairs spanning 35,000 distinct prompts. Leveraging this open resource, named Pick-a-Pic, the authors fine-tuned an image-text scoring model called PickScore to estimate how satisfied a user is with a generated image relative to their prompt.
Key findings demonstrate the superiority of authentic user data over conventional metrics. First, PickScore achieved a 70.5% accuracy rate in predicting user choices on held-out test data, outperforming both human expert annotators (68.0%) and existing baselines like standard CLIP-H (60.8%) and aesthetic predictors (56.8%). Second, traditional automated benchmarks were shown to be deeply flawed; for example, the widely used Fréchet Inception Distance metric exhibited a strong negative correlation (-0.900) with human rankings on standard image captions, whereas PickScore showed a strong positive correlation (0.917). Third, when evaluating models against actual user preferences, PickScore's rankings achieved a 0.790 correlation with ground truth, significantly outperforming alternative scoring metrics. Finally, using PickScore to automatically select the best image out of 100 generated variations yielded human preference win rates exceeding 71% against unranked model outputs and competing scoring functions.
These results carry significant strategic implications for machine learning development. Traditional evaluation pipelines rely on photographic caption datasets like MS-COCO, which do not reflect the creative, fictional prompts real users actually write. Furthermore, optimizing models against outdated realism metrics can inadvertently degrade the vivid, aesthetically pleasing qualities users prefer. Incorporating a robust automated judge like PickScore provides an efficient, low-cost way to boost generation quality through post-processing rank selection or direct model alignment (such as reinforcement learning from human feedback) without requiring continuous human labeling.
Organizations developing or deploying text-to-image systems should immediately replace or supplement standard benchmarks with Pick-a-Pic prompts and adopt PickScore for automated performance assessments. When deploying generative applications, engineering teams should evaluate using PickScore as a reranking filter over candidate images to improve output quality. Decision-makers should note that while the dataset underwent strict filtering, some residual not-safe-for-work content and user bias may remain. Nonetheless, the high statistical consistency across extensive trials provides strong confidence in PickScore as a primary evaluation and selection tool.
- Paper: CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Jack Hessel et al. (2021). This paper introduces CLIPScore for evaluating image-text alignment using CLIP embeddings, establishing the foundational metric and paradigm that PickScore directly builds upon and outperforms.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. This work introduces CLIP, the core vision-language architecture and pretrained representation used to train the PickScore preference scoring function.
- Paper: Exploring CLIP for Assessing the Look and Feel of Images, Jianyi Wang et al. (2022). This study explores using CLIP embeddings for assessing aesthetic and technical image quality, providing important groundwork for training CLIP-based models on human visual judgments.
- Paper: Microsoft COCO Captions: Data Collection and Evaluation Server, Xinlei Chen et al. (2015). This paper presents the MS-COCO Captions dataset, which serves as the traditional text-to-image benchmark that Pick-a-Pic seeks to improve upon and replace with more realistic user prompts.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). This work introduces photorealistic text-to-image diffusion models along with DrawBench for evaluating text-image alignment, contextualizing the generation architectures and evaluation challenges addressed by Pick-a-Pic.
- Paper: LAION-5B: An open large-scale dataset for training next generation image-text models, Christoph Schuhmann et al. (2022). This paper introduces the open-access LAION-5B dataset, providing key background on community-driven open data initiatives for training large-scale text-to-image models.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). This paper develops Stable Diffusion XL (SDXL), advancing the family of open text-to-image foundation models evaluated and enhanced by human-preference datasets and scoring metrics like PickScore.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This work scales rectified flow transformers for high-resolution image synthesis, utilizing human-preference alignments and advanced benchmarking principles relevant to PickScore.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). This paper introduces Chatbot Arena, an open web platform that generalizes crowdsourced pairwise human preference collection and live Elo evaluation to conversational language models.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). This study develops SpatialGenEval to systematically benchmark spatial reasoning in text-to-image generation, building on the broader movement toward richer, prompt-based evaluation protocols.
