GenEval: An object-focused framework for evaluating text-to-image alignment

Dhruba GhoshHannaneh HajishirziLudwig Schmidt

article2023NeurIPS693 citations

Introduces GenEval, an automated evaluation framework that uses object detection models to evaluate fine-grained compositional properties in text-to-image generation—such as object counts, spatial relations, and color binding—with strong human agreement.

Listen

Recent rapid progress in artificial intelligence has generated thousands of text-to-image generative models. However, standard automated evaluation metrics assess only holistic image realism or general image-text similarity without verifying specific compositional requirements in text prompts. Because manual human evaluation is too slow and expensive to keep pace with model development, automated, fine-grained benchmarking tools are urgently needed to assess whether models follow detailed instructions.

The article demonstrates an automated, object-focused evaluation framework called GENEVAL. The objective is to provide a reliable, modular, and interpretable method for evaluating whether text-to-image models accurately render specific objects, exact counts, spatial arrangements, and associated visual attributes like color.

To accomplish this, the authors linked existing discriminative vision models together without requiring specialized training or synthetic data. The pipeline uses an object detection and instance segmentation model to verify object presence, count, and relative positioning from image bounding boxes. It then crops detected objects, masks the background, and applies a zero-shot image classifier to verify colors. The framework was validated through a human study involving 6,000 fine-grained annotations across 1,200 generated images, comparing its judgment against human consensus and standard automated metrics. The authors then benchmarked several major open-source text-to-image models across 553 standardized prompts covering six compositional tasks.

The findings establish that GENEVAL closely mirrors human perception while exposing critical performance gaps across current generative models. First, the automated framework achieves 83% overall agreement with human annotators—closely approaching the 88% inter-annotator agreement rate—and outscores standard similarity metrics like CLIPScore on complex compositional tasks (improving human agreement on counting tasks by about 22 percentage points). Second, the DeepFloyd IF-XL model achieved the highest overall score at 61% accuracy, outperforming Stable Diffusion v2.1 (50%) and Stable Diffusion XL (55%). Third, while models handle single objects (97–98% accuracy) and single colors (81–85% accuracy) effectively, all models perform poorly on complex spatial arrangements and attribute binding; even the best systems reached only 15% accuracy for relative positioning and 35% for binding specific colors to multiple objects. Fourth, increasing vision model size consistently improved performance on multi-object rendering and color binding, but scaling model parameters did not resolve spatial positioning failures, and extending pretraining iterations without architectural upgrades yielded negligible benefits.

These findings mean that text-to-image systems remain unreliable for production applications requiring precise spatial relationships or exact object counts. For organizations deploying these tools, holistic quality scores can create a false sense of accuracy, whereas instance-level breakdown helps developers diagnose concrete model flaws—such as directional positioning biases and color leakage. The modular nature of the evaluation pipeline allows organizations to upgrade individual vision components as better detection systems emerge, reducing evaluation costs and development cycle times.

The article recommends that model developers focus future research on improving text encoders and dataset composition rather than simply extending pretraining runs or expanding image-generation parameters alone. Developers should also utilize granular diagnostic evaluations to identify and fix specific failure patterns early. Because current object detectors are bound to common photographic datasets, users should exercise caution when evaluating non-photographic art styles, hand anatomy, or niche object classes that fall outside standard object taxonomies until open-vocabulary detectors are fully integrated.

Cover for GenEval: An object-focused framework for evaluating text-to-image alignment

Abstract

Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are critical for evaluating the increasingly large number of new models. However, most current automated evaluation metrics like FID or CLIPScore only offer a holistic measure of image quality or image-text alignment, and are unsuited for fine-grained or instance-level analysis. In this paper, we introduce GenEval, an object-focused framework to evaluate compositional image properties such as object co-occurrence, position, count, and color. We show that current object detection models can be leveraged to evaluate text-to-image models on a variety of generation tasks with strong human agreement, and that other discriminative vision models can be linked to this pipeline to further verify properties like object color. We then evaluate several open-source text-to-image models and analyze their relative generative capabilities on our benchmark. We find that recent models demonstrate significant improvement on these tasks, though they are still lacking in complex capabilities such as spatial relations and attribute binding. Finally, we demonstrate how GenEval might be used to help discover existing failure modes, in order to inform development of the next generation of text-to-image models. Our code to run the GenEval framework is publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 GenEval: Our object-focused evaluation framework
  • 3.1 Setup
  • 3.2 Evaluation framework
  • 4 Measuring alignment with human judgment
  • 5 Benchmarking progress in recent T2I models
  • 5.1 Experiments
  • 5.2 Results
  • 6 Limitations
  • 7 Conclusion
  • References
  • A Further experiments
  • A.1 Alignment with human judgment
  • A.2 Evaluation parameters
  • A.3 Qualitative examples
  • B Ethics statement
  • C Technical details
  • C.1 Prompt generation
  • C.2 Image generation
  • C.3 Evaluation
  • D Human study details
  • D.1 Full text

Knowls

  1. Knowl 1 — GENEVAL decomposes prompt compliance into interpretable object-level checks

    model/method

    GENEVAL evaluates a text-to-image image against the objects and properties specified in its prompt instead of assigning only a holistic image-quality or image–text-alignment score. It uses object detection and instance segmentation to check object presence and obtain per-instance locations and masks; those outputs support checks of object count and relative position, and can be passed to separate discriminative vision models to assess attributes such as color. Each image receives a binary correctness judgment: it is correct only when all prompt-specified elements checked by the framework are satisfied. For an incorrect image, the framework reports which expected objects or properties failed, making the score interpretable at the instance level.

  2. Knowl 2 — GENEVAL combines detection, geometry, and zero-shot color classification

    model/method

    The evaluated implementation uses a Mask2Former instance-segmentation model with a Swin-S backbone, trained on MS COCO. It retains detections above confidence 0.3 for ordinary tasks and raises the threshold to 0.9 for counting, where low-confidence detections can inflate instance counts. Object presence is checked first; counting additionally requires the detected count to match the prompt, while extra instances are not penalized in other tasks. Relative position is estimated from bounding-box centroids, with a minimum visible-offset threshold scaled to the boxes’ dimensions; the selected threshold is c=0.1c=0.1. For color and attribute-binding tasks, each detected object is cropped to its bounding box, pixels outside its segmentation mask are replaced with gray, and CLIP ViT-L/14 performs zero-shot classification among color–object text prompts. The image-level score is 1 only if all requested checks pass and 0 otherwise; task scores average image scores, and the overall score averages the six task scores.

  3. Knowl 3 — The benchmark covers six compositional generation tasks

    experimental setup

    GENEVAL evaluates single-object generation, generation of two different objects, counting, object color, relative position, and attribute binding (two objects with distinct specified colors). Its templates use 80 MS COCO object categories; the benchmark contains 80 single-object prompts, 99 two-object prompts, 80 counting prompts, 94 color prompts, 100 position prompts, and 100 attribute-binding prompts, for 553 prompts total. Counting requests two, three, or four instances; position uses above, below, left of, or right of. The color-related prompts use basic color terms, and gray is excluded from generated color choices. For each prompt, the experiments generate four images (or retrieve four images for the real-image baseline), using each model’s default generation settings.

  4. Knowl 4 — Human evaluation supports GENEVAL’s alignment with annotators

    empirical result

    The human evaluation covered 1,200 images: 400 each from Stable Diffusion v2.1, IF-XL, and LAION-5B CLIP retrieval. Five annotators assessed each image, yielding 6,000 annotations; they identified visible objects and judged prompt-relevant properties including count, color, realism, relative position, and overall caption fit. Across images, GENEVAL achieved 83% agreement with human judgments, compared with 80% for CLIPScore and 88% pairwise interannotator agreement. On the 860 images for which annotators were unanimous, GENEVAL agreement was 91%, versus 87% for CLIPScore. With CLIPScore thresholds tuned separately by task, GENEVAL agreed more often on the more compositional tasks, including counting, position, and attribute binding; CLIPScore was slightly stronger on simpler single-object and color tasks.

  5. Knowl 5 — IF-XL leads the evaluated open-source models, while difficult compositions remain weak

    empirical result

    The benchmark compares models on 553 prompts with four generated or retrieved images per prompt. Task and overall GENEVAL scores are proportions; CLIPScore is reported on a 0–100 scale. The human column gives human scores where reported. IF-XL has the highest overall GENEVAL score, 0.61, followed by SD-XL at 0.55 and SDv2.1 at 0.50. Position and attribute binding remain difficult even for the strongest models.

    ModelSingle objectTwo objectsCountingColorsPositionAttribute bindingOverallCLIPScoreHuman
    CLIP retrieval0.890.220.370.620.030.000.3527.80.42
    minDALL-E0.730.110.120.370.020.010.2327.3—
    SDv1.50.970.380.350.760.040.060.4333.5—
    SDv2.10.980.510.440.850.070.170.5036.20.57
    SD-XL0.980.740.390.850.150.230.5536.7—
    IF-XL0.970.740.660.810.130.350.6136.50.72

    Single-object and color tasks are relatively strong across models, whereas position and attribute-binding scores are low. IF-XL leads in counting and attribute binding; SD-XL and IF-XL both score 0.74 on two-object generation. The CLIP-retrieval baseline performs poorly overall but exceeds minDALL-E and SDv1.5 on counting. Overall scores vary by approximately 0.01–0.02 across random seeds.

  6. Knowl 6 — Increasing IF model scale helps some capabilities, but more Stable Diffusion training does not reliably help

    empirical result

    Across IF-M, IF-L, and IF-XL, all of which use the T5-XXL text encoder, larger models improve two-object generation, counting, and especially attribute binding; position does not improve monotonically. The paper attributes the scale change to increasingly sized image-generation and upscaling modules. By contrast, the Stable Diffusion v1.1–v1.5 checkpoints, trained with increasing amounts of LAION-5B data, show no consistent overall gain after v1.2. Scores rise from 0.44 for v1.2 to 0.51 for v2.0, but that comparison also changes the text encoder and other training details, so the paper does not isolate a single cause. Scores below are overall, two-object, counting, position, and attribute binding, in that order.

    ModelScores
    IF-M0.52, 0.53, 0.53, 0.10, 0.14
    IF-L0.54, 0.60, 0.57, 0.15, 0.17
    IF-XL0.61, 0.74, 0.66, 0.13, 0.35
    SD v1.10.41, 0.31, 0.33, 0.02, 0.05
    SD v1.20.44, 0.41, 0.37, 0.03, 0.10
    SD v1.30.43, 0.38, 0.35, 0.03, 0.05
    SD v1.40.42, 0.36, 0.35, 0.01, 0.07
    SD v1.50.43, 0.38, 0.35, 0.04, 0.06
    SD v2.00.51, 0.50, 0.48, 0.06, 0.15
    SD v2.10.50, 0.51, 0.44, 0.07, 0.17
  7. Knowl 7 — Fine-grained scores expose systematic generation errors

    empirical result

    GENEVAL’s per-property outputs reveal errors that a single alignment score would not localize. In the position task, IF-XL shows a bias: when it generates both requested objects, the first-mentioned object is more often left of the second than right of it, even though prompts are evenly distributed among the four requested directions. In attribute binding, Stable Diffusion v2.1 is more prone than IF-XL to swapping the two specified colors between objects. The paper also identifies color leakage onto background objects as a failure pattern. These analyses illustrate how GENEVAL can diagnose directional and attribute-binding errors, rather than only ranking models.

  8. Knowl 8 — Task-specific thresholds improve counting and position judgments

    empirical result

    The authors tuned evaluation parameters against human annotations because default detector outputs can misrepresent generated images. For counting, detections at the default 0.3 confidence threshold often include excess low-confidence boxes; raising the threshold to 0.9 increases Cohen’s kappa agreement from 0.37 to 0.65. Five-fold cross-validation reports 0.823 ± 0.013 agreement on validation splits for the selected counting threshold. For position, a small minimum centroid offset prevents overlapping or nearly touching objects from being classified as visibly separated; the chosen scale factor is c=0.1c=0.1, and five-fold validation reports 0.822 ± 0.013 agreement. Non-maximum suppression did not improve detector performance in the counting cases tested.

  9. Knowl 9 — Cropping objects and masking their backgrounds improves color judgments

    empirical result

    The authors measure color-classification preprocessing by Cohen’s kappa agreement with human annotators. Cropping a detected object to its bounding box and replacing pixels outside its segmentation mask with gray each help, particularly for attribute binding, where other image objects can distract the classifier. Using both steps gives the strongest GENEVAL agreement among the tested preprocessing options, although masking alone is close for attribute binding.

    MethodCropBackground maskColors κAttribute-binding κ
    GENEVALNoNo0.320.01
    GENEVALYesNo0.370.33
    GENEVALNoYes0.430.47
    GENEVALYesYes0.450.49
    CLIPScore——0.230.47
    Interannotator——0.550.65
  10. Knowl 10 — GENEVAL is constrained by detector coverage and out-of-distribution images

    limitation

    GENEVAL’s reliability and scope depend on its object detector and the available downstream discriminative models. The evaluated detector is trained on MS COCO, whose finite class set and category granularity limit what can be checked: for example, the framework can count detected people but cannot assess the number of fingers on their hands. Detectors trained primarily on photographs can also fail on visually different inputs such as clip art, and overlapping same-class objects can be merged or otherwise misdetected. The paper notes that stronger open-vocabulary detectors trained on broader image distributions could reduce these constraints, but presents this as a prospective possibility rather than a demonstrated result.

Coverage note — Qualitative image galleries and the detailed annotation questionnaire are omitted because they add examples and procedural detail, but no substantial contributed method, result, analysis, or stated limitation is deliberately left out.

References

  1. 1.URL https://openai.com/dall-e-2.
  2. 2.URL https://www.midjourney.com/.
  3. 3.Romain Beaumont. Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them. https://github.com/rom1504/clip-retrieval, 2022.
  4. 4.Brent Berlin and Paul Kay. Basic color terms: Their universality and evolution. CLSI Publ., 2000.
  5. 5.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  6. 6.Yixiong Chen, Li Liu, and Chris Ding. X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models, 2023.
  7. 7.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. 2022.
  8. 8.Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2818–2829, 2023.
  9. 9.Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers. 2022.
  10. 10.Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Phuc Le Khac, Luke Melas, and Ritobrata Ghosh. Dall·e mini, 7 2021. URL https://github.com/borisdayma/dalle-mini.
  11. 11.Deep-Floyd. Deep-floyd/if. URL https://github.com/deep-floyd/IF.
  12. 12.Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. arXiv preprint arXiv:2303.11331, 2023.
  13. 13.Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Reddy Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured diffusion guidance for compositional text-to-image synthesis. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=PUIqjT4rzq7.
  14. 14.Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vineet, Eric Horvitz, Ece Kamar, Chitta Baral, and Yezhou Yang. Benchmarking spatial relationships in text-to-image generation, 2022.
  15. 15.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: a reference-free evaluation metric for image captioning. In EMNLP, 2021.
  16. 16.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, pp. 6629–6640, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  17. 17.Tobias Hinz, Stefan Heinrich, and Stefan Wermter. Semantic object accuracy for generative text-to-image synthesis. IEEE transactions on pattern analysis and machine intelligence, 44(3):1552–1565, 2020.
  18. 18.Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee. Inferring semantic layout for hierarchical text-to-image synthesis. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7986–7994, 2018.
  19. 19.Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. arXiv preprint arXiv:2303.11897, 2023.
  20. 20.Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park, Eli Shechtman, Sylvain Paris, and Taesung Park. Scaling up gans for text-to-image synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  21. 21.Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation, 2023.
  22. 22.Kuprel. Kuprel/min-dalle: Min(dall·e) is a fast, minimal port of dall·e mini to pytorch. URL https://github.com/kuprel/min-dalle.
  23. 23.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll’a r, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. CoRR, abs/1405.0312, 2014. URL http://arxiv.org/abs/1405.0312.
  24. 24.Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation, 2023.
  25. 25.Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov. Generating images from captions with attention. In ICLR, 2016.
  26. 26.Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, et al. Simple open-vocabulary object detection with vision transformers. arXiv preprint arXiv:2205.06230, 2022.
  27. 27.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  28. 28.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL ’02, pp. 311–318, USA, 2002. Association for Computational Linguistics. doi: 10.3115/1073083.1073135. URL https://doi.org/10.3115/1073083.1073135.
  29. 29.Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. Benchmark for compositional text-to-image synthesis. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021. URL https://openreview.net/forum?id=bKBhQhPeKaF.
  30. 30.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.
  31. 31.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  33. 33.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 8821–8831. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/ramesh21a.html.
  34. 34.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  35. 35.Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis. CoRR, abs/1605.05396, 2016. URL http://arxiv.org/abs/1605.05396.
  36. 36.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  37. 37.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding, 2022.
  38. 38.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, pp. 2234–2242, Red Hook, NY, USA, 2016. Curran Associates Inc. ISBN 9781510838819.
  39. 39.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5b: An open large-scale dataset for training next generation image-text models. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview.net/forum?id=M3Y74vmsMcY.
  40. 40.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. CoRR, abs/1411.5726, 2014.
  41. 41.Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis, 2023.
  42. 42.Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagereward: Learning and evaluating human preferences for text-to-image generation, 2023.
  43. 43.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. CoRR, abs/1711.10485, 2017.
  44. 44.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=AFDcYJKhND. Featured Certification.

Citation

MLA
Ghosh, D., et al. “GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment”. arXiv, 2023, http://arxiv.org/abs/2310.11513v1.
APA
Ghosh, D., Hajishirzi, H., & Schmidt, L. (2023). GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. arXiv. http://arxiv.org/abs/2310.11513v1
Chicago
Ghosh, D., H. Hajishirzi, and L. Schmidt. 2023. “GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment”. arXiv. http://arxiv.org/abs/2310.11513v1.
Harvard
Ghosh, D., Hajishirzi, H. and Schmidt, L. (2023) “GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.11513v1.
Vancouver
1. Ghosh D, Hajishirzi H, Schmidt L (2023) GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. arXiv

BibTeX

@article{ghosh2023geneval,
  title = {GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment},
  author = {Ghosh, Dhruba and Hajishirzi, Hanna and Schmidt, Ludwig},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.11513v1},
  eprint = {2310.11513}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors