GenEval: An object-focused framework for evaluating text-to-image alignment
Dhruba GhoshHannaneh HajishirziLudwig Schmidt
Introduces GenEval, an automated evaluation framework that uses object detection models to evaluate fine-grained compositional properties in text-to-image generation—such as object counts, spatial relations, and color binding—with strong human agreement.
Recent rapid progress in artificial intelligence has generated thousands of text-to-image generative models. However, standard automated evaluation metrics assess only holistic image realism or general image-text similarity without verifying specific compositional requirements in text prompts. Because manual human evaluation is too slow and expensive to keep pace with model development, automated, fine-grained benchmarking tools are urgently needed to assess whether models follow detailed instructions.
The article demonstrates an automated, object-focused evaluation framework called GENEVAL. The objective is to provide a reliable, modular, and interpretable method for evaluating whether text-to-image models accurately render specific objects, exact counts, spatial arrangements, and associated visual attributes like color.
To accomplish this, the authors linked existing discriminative vision models together without requiring specialized training or synthetic data. The pipeline uses an object detection and instance segmentation model to verify object presence, count, and relative positioning from image bounding boxes. It then crops detected objects, masks the background, and applies a zero-shot image classifier to verify colors. The framework was validated through a human study involving 6,000 fine-grained annotations across 1,200 generated images, comparing its judgment against human consensus and standard automated metrics. The authors then benchmarked several major open-source text-to-image models across 553 standardized prompts covering six compositional tasks.
The findings establish that GENEVAL closely mirrors human perception while exposing critical performance gaps across current generative models. First, the automated framework achieves 83% overall agreement with human annotators—closely approaching the 88% inter-annotator agreement rate—and outscores standard similarity metrics like CLIPScore on complex compositional tasks (improving human agreement on counting tasks by about 22 percentage points). Second, the DeepFloyd IF-XL model achieved the highest overall score at 61% accuracy, outperforming Stable Diffusion v2.1 (50%) and Stable Diffusion XL (55%). Third, while models handle single objects (97–98% accuracy) and single colors (81–85% accuracy) effectively, all models perform poorly on complex spatial arrangements and attribute binding; even the best systems reached only 15% accuracy for relative positioning and 35% for binding specific colors to multiple objects. Fourth, increasing vision model size consistently improved performance on multi-object rendering and color binding, but scaling model parameters did not resolve spatial positioning failures, and extending pretraining iterations without architectural upgrades yielded negligible benefits.
These findings mean that text-to-image systems remain unreliable for production applications requiring precise spatial relationships or exact object counts. For organizations deploying these tools, holistic quality scores can create a false sense of accuracy, whereas instance-level breakdown helps developers diagnose concrete model flaws—such as directional positioning biases and color leakage. The modular nature of the evaluation pipeline allows organizations to upgrade individual vision components as better detection systems emerge, reducing evaluation costs and development cycle times.
The article recommends that model developers focus future research on improving text encoders and dataset composition rather than simply extending pretraining runs or expanding image-generation parameters alone. Developers should also utilize granular diagnostic evaluations to identify and fix specific failure patterns early. Because current object detectors are bound to common photographic datasets, users should exercise caution when evaluating non-photographic art styles, hand anatomy, or niche object classes that fall outside standard object taxonomies until open-vocabulary detectors are fully integrated.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Imagen’s DrawBench established challenging prompts for testing counting, color, and other text-to-image alignment failures that GENEVAL evaluates at the object level.
- Paper: When and why vision-language models behave like bags-of-words, and what to do about it?, Mert Yuksekgonul et al. (2022). ARO’s tests of attribute binding and object relations introduce the compositional weaknesses that GENEVAL measures through object-level attributes, counts, and positions.
- Paper: VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena, Letitia Parcalabescu et al. (2022). VALSE’s diagnostic tasks for counting and spatial relations provide a foundation for understanding GENEVAL’s fine-grained evaluation of prompt following.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). SpatialGenEval takes the spatial failures exposed by object-focused evaluation into more detailed scenes, testing localization, spatial reasoning, and interactions.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun et al. (2025). T2V-CompBench carries compositional alignment evaluation into video, extending object, attribute, and spatial tests to temporal relationships and motion.
