T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

Kaiyi HuangKaiyue SunEnze XieZhenguo LiXihui Liu

article2023NeurIPS280 citations

Proposes a 6,000-prompt benchmark with human-correlated evaluation metrics to systematically assess and improve attribute binding and object relationships in compositional text-to-image generation.

Listen

Modern artificial intelligence models can create realistic images from written descriptions, but they frequently fail when asked to combine multiple concepts accurately into a single scene. For example, when prompted to depict a blue bench next to a red car, these systems often mix up the colors or misplace the objects entirely. This limitation poses substantial challenges for deploying generative tools in practical visual workflows, marketing, and design, where precise adherence to instructions is critical.

To address this gap, the article introduces a standardized benchmark named T2I-CompBench, which establishes a framework for evaluating and enhancing compositional image generation across 6,000 text prompts. The benchmark divides compositional challenges into three primary categories: binding attributes (such as color, shape, and texture) to the correct objects, establishing relationships (both spatial positioning and actions) between objects, and handling complex scenes involving multiple objects and mixed descriptors. In parallel, the article develops specialized evaluation metrics—including a visual question-answering method that breaks prompts into single object-attribute queries, an object-detection metric for spatial layouts, and a composite score for complex prompts—alongside an efficient fine-tuning method named GORS that trains models using only high-quality, well-aligned generated images.

Key findings show that existing evaluation tools, which rely on general image-text similarity, fail to capture fine-grained layout and attribute errors, whereas the proposed targeted metrics correlate significantly better with human judgment. Among all evaluated generation tasks, spatial positioning proved to be the most difficult for current models, yielding low overall alignment scores, whereas non-spatial interactions were the easiest. Across six benchmarked generative models, the proposed GORS fine-tuning strategy consistently outperformed prior systems across all categories, achieving top human alignment scores such as 0.83 on color binding and 0.46 on spatial relationships, while scaling predictably as training data increased.

These findings indicate that current foundational models cannot be assumed to understand complex multi-object prompts out of the box, requiring organizations to implement targeted training and domain-specific validation rather than relying on standard similarity metrics. By showing that models can be effectively aligned using reward-weighted fine-tuning on high-performing generated samples, the results suggest a cost-effective pathway to improve visual reliability without massive re-training from scratch.

Decision-makers and engineering teams looking to adopt text-to-image systems should incorporate modular, specialized evaluation pipelines instead of single generalist metrics to monitor performance. Furthermore, adopting reward-filtered fine-tuning offers a viable option to boost generation accuracy, with performance gains scaling alongside the volume of training prompts.

The study notes several limitations, including the lack of a single unified evaluation metric across all compositional tasks and a restriction to two-dimensional spatial layouts rather than full three-dimensional spatial reasoning. Current large multimodal vision-language evaluators also exhibit limitations such as hallucinations and visual misinterpretations. While confidence in the benchmark and targeted metric correlations is high, practitioners should exercise caution regarding potential biases and hallucinations inherent in automated evaluation tools.

Cover for T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

Abstract

Despite the stunning ability to generate high-quality images by recent text-to-image models, current approaches often struggle to effectively compose objects with different attributes and relationships into a complex and coherent scene. We propose T2I-CompBench, a comprehensive benchmark for open-world compositional text-to-image generation, consisting of 6,000 compositional text prompts from 3 categories (attribute binding, object relationships, and complex compositions) and 6 sub-categories (color binding, shape binding, texture binding, spatial relationships, non-spatial relationships, and complex compositions). We further propose several evaluation metrics specifically designed to evaluate compositional text-to-image generation and explore the potential and limitations of multimodal LLMs for evaluation. We introduce a new approach, Generative mOdel finetuning with Reward-driven Sample selection (GORS), to boost the compositional text-to-image generation abilities of pretrained text-to-image models. Extensive experiments and evaluations are conducted to benchmark previous methods on T2I-CompBench, and to validate the effectiveness of our proposed evaluation metrics and GORS approach. Project page is available at https://karine-h.github.io/T2I-CompBench/.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 T2I-CompBench
  • 3.1 Attribute Binding
  • 3.2 Object Relationship
  • 3.3 Complex Compositions
  • 4 Evaluation Metrics
  • 4.1 Disentangled BLIP-VQA for Attribute Binding Evaluation
  • 4.2 UniDet-based Spatial Relationship Evaluation
  • 4.3 3-in-1 Metric for Complex Compositions Evaluation
  • 4.4 Evaluation with Multimodal Large Language Models
  • 5 Method
  • 6 Experiments
  • 6.1 Experimental Setup
  • 6.2 Evaluation Metrics
  • 6.3 Quantitative and Qualitative Evaluation
  • 6.4 Human Correlation of the Evaluation Metrics
  • 6.5 Ablation study
  • 7 Conclusion and Discussions
  • Acknowledgements
  • References
  • A Implementation Details
  • B T2I-CompBench Dataset Construction
  • C Evaluation Metrics
  • C.1 Prompts for MiniGPT4-CoT and MiniGPT4 Evaluation
  • C.2 Human Evaluation
  • D Additional Results
  • D.1 Quantitative Results of Seen and Unseen Splits
  • D.2 MiniGPT-4 Evaluation without Chain-of-Thought
  • D.3 Reward models to Select Samples for GORS-unbiased
  • D.4 Scalability of our proposed approach
  • D.5 Qualitative Results of Ablation Study
  • D.6 Qualitative Results and Comparison with Prior Work
  • E Limitation and Potential Negative Social Impacts

Knowls

  1. Knowl 1 — T2I-CompBench Benchmark Specification

    definition

    T2I-CompBench is a benchmark designed to evaluate open-world compositional text-to-image generation. The benchmark comprises 6,000 compositional text prompts organized into 3 categories and 6 sub-categories:

    1. Attribute Binding (3,000 prompts): Requires correctly binding visual attributes to multiple distinct objects in the scene.

      • Color Binding (1,000 prompts): Prompts specifying colors for at least two objects (e.g., "a red flower and a yellow vase").
      • Shape Binding (1,000 prompts): Prompts specifying geometric shapes for at least two objects across 21 shape attributes (e.g., circular, cubic, conical, diamond).
      • Texture Binding (1,000 prompts): Prompts specifying materials and textures for at least two objects (e.g., rubber, plastic, metallic, wooden, fabric, fluffy, leather, glass). Within each attribute binding sub-category, 800 prompts follow the template "a {adj} {noun} and a {adj} {noun}" and 200 prompts use natural phrasing without a predefined template.
    2. Object Relationships (2,000 prompts): Evaluates interactions and layouts between at least two objects.

      • Spatial Relationships (1,000 prompts): Prompts using 7 spatial prepositions: "on the side of", "next to", "near", "on the left of", "on the right of", "on the bottom of", and "on the top of". Contrastive prompts are created by swapping the order of the two nouns.
      • Non-Spatial Relationships (1,000 prompts): Prompts capturing actions and interactions (e.g., "watch", "wear", "hold", "play with", "walk with", "sit on").
    3. Complex Compositions (1,000 prompts): Evaluates complex open-world scenes partitioned equally into four scenarios (250 prompts each):

      • Two objects with multiple attributes.
      • Two objects with mixed attribute types (e.g., shape and color).
      • More than two objects with multiple attributes.
      • More than two objects with mixed attribute types.

    For each sub-category, the 1,000 prompts are split into 700 training prompts and 300 testing prompts. For attribute binding test sets, prompts are further partitioned into 200 seen attribute-noun pairs and 100 unseen attribute-noun pairs. The benchmark vocabulary contains 2,316 nouns, 33 colors, 32 shapes, 23 textures, 7 spatial relationships, and 875 non-spatial relationships.

  2. Knowl 2 — Disentangled BLIP-VQA Evaluation Metric for Attribute Binding

    model/method

    Disentangled BLIP-VQA is an evaluation metric designed to assess attribute binding in text-to-image generation. Holistic captioning models or standard visual question answering (VQA) often confuse multiple object-attribute pairings or fail to report granular attributes.

    Given a text prompt containing multiple object-attribute pairs (such as "a green bench and a red car") and a synthesized image xx, the metric decomposes the prompt into MM independent, single-concept questions:

    • Question 1: "a green bench?"
    • Question 2: "a red car?"

    Each question qiq_i (i∈{1,…,M}i \in \{1, \dots, M\}) is provided alongside image xx to a pretrained BLIP VQA model (BLIP with ViT-B and CapFilt-L). The model outputs the probability of answering "yes", denoted as P("yes"∣x,qi)∈[0,1]P(\text{"yes"} \mid x, q_i) \in [0, 1].

    The overall attribute binding score SB-VQA(x)S_{\text{B-VQA}}(x) is the product of the probabilities across all decomposed questions: SB-VQA(x)=∏i=1MP("yes"∣x,qi)S_{\text{B-VQA}}(x) = \prod_{i=1}^{M} P(\text{"yes"} \mid x, q_i)

    This multiplicative formulation requires the generative model to successfully bind all specified attributes simultaneously, assigning a low score if any individual attribute-object pairing is missing or bound incorrectly.

  3. Knowl 3 — UniDet-Based Evaluation Metric for 2D Spatial Relationships

    model/method

    The UniDet-based evaluation metric evaluates 2D spatial relationships between pairs of objects in generated images using the UniDet object detector trained across COCO, Objects365, OpenImages, and Mapillary.

    Given a generated image and a prompt specifying a spatial relationship between object 1 and object 2, UniDet detects bounding boxes for both objects. Let the center coordinates of the detected boxes be (x1,y1)(x_1, y_1) for object 1 and (x2,y2)(x_2, y_2) for object 2, with bounding box Intersection-over-Union denoted as IoU\text{IoU}.

    The evaluation rules are defined as follows:

    • Left / Right / Top / Bottom:
      • Left ("on the left of"): Satisfied if x1<x2x_1 < x_2, ∣x1−x2∣>∣y1−y2∣|x_1 - x_2| > |y_1 - y_2|, and IoU<0.1\text{IoU} < 0.1.
      • Right ("on the right of"): Satisfied if x1>x2x_1 > x_2, ∣x1−x2∣>∣y1−y2∣|x_1 - x_2| > |y_1 - y_2|, and IoU<0.1\text{IoU} < 0.1.
      • Top ("on the top of"): Satisfied if y1<y2y_1 < y_2, ∣y1−y2∣>∣x1−x2∣|y_1 - y_2| > |x_1 - x_2|, and IoU<0.1\text{IoU} < 0.1.
      • Bottom ("on the bottom of"): Satisfied if y1>y2y_1 > y_2, ∣y1−y2∣>∣x1−x2∣|y_1 - y_2| > |x_1 - x_2|, and IoU<0.1\text{IoU} < 0.1.
    • Proximity ("next to", "near", "on the side of"): Satisfied if the Euclidean distance (x1−x2)2+(y1−y2)2\sqrt{(x_1 - x_2)^2 + (y_1 - y_2)^2} is below a predefined distance threshold.

    If the detected bounding boxes satisfy the target condition, the metric assigns an alignment score of 1; otherwise, it assigns 0.

  4. Knowl 4 — 3-in-1 Metric and Chain-of-Thought MLLM Evaluation for Compositional Generation

    model/method

    To evaluate complex compositions featuring multiple objects, mixed attribute types, and spatial/non-spatial relationships simultaneously, two evaluation frameworks are introduced:

    1. 3-in-1 Composite Metric: Because individual metrics specialize in distinct compositionality types (Disentangled BLIP-VQA for attributes, UniDet for spatial layouts, and CLIPScore for non-spatial semantic interactions), the 3-in-1 metric calculates the arithmetic mean of all three scores: S3-in-1=SCLIP+SB-VQA+SUniDet3S_{\text{3-in-1}} = \frac{S_{\text{CLIP}} + S_{\text{B-VQA}} + S_{\text{UniDet}}}{3}

    2. MiniGPT4-Chain-of-Thought (mGPT-CoT) Evaluation: A multimodal large language model (MiniGPT-4 with Vicuna 13B) evaluates alignment via a two-stage sequential prompting strategy:

      • Stage 1 (Image Description): The model is prompted to identify all objects, attributes, and relationships within a 50-word budget.
      • Stage 2 (Alignment Scoring): Conditioned on the image and its own description from Stage 1, the model rates alignment from 0 to 100 according to a structured rubric:
        • 100: Image perfectly matches the content with no discrepancies.
        • 80: Most actions, events, and relationships portrayed with minor discrepancies.
        • 60: Some elements depicted, but key parts or details omitted/incorrect.
        • 40: Image fails to convey the scope of the prompt.
        • 20: Image almost irrelevant to the prompt.
  5. Knowl 5 — Generative Model Finetuning with Reward-Driven Sample Selection (GORS)

    model/method

    Generative mOdel finetuning with Reward-driven Sample selection (GORS) is a method to enhance the compositional generation capabilities of pretrained diffusion models by fine-tuning on self-generated samples weighted by alignment rewards.

    Given a pretrained text-to-image diffusion model pθp_\theta and a set of compositional text prompts {y1,y2,…,yn}\{y_1, y_2, \dots, y_n\}:

    1. Sample Generation: Generate kk candidate images for each prompt yiy_i, yielding k⋅nk \cdot n images {x1,x2,…,xkn}\{x_1, x_2, \dots, x_{kn}\}.
    2. Reward Scoring and Thresholding: Predict text-image alignment scores sj∈[0,1]s_j \in [0, 1] for each generated sample (xj,y)(x_j, y) using an alignment reward model. A threshold τ\tau is applied to select high-alignment samples: Ds={(xj,y,sj)∣sj>τ}\mathcal{D}_s = \{(x_j, y, s_j) \mid s_j > \tau\}
    3. Reward-Weighted Parameter-Efficient Finetuning: Fine-tune both the CLIP text encoder self-attention layers and the U-Net attention layers using Low-Rank Adaptation (LoRA).

    To prevent metric-gaming bias during benchmarking, an unbiased variant (GORS-unbiased) adopts reward models entirely distinct from the evaluation metrics:

    • Attribute Binding Reward: Grounded-SAM extracts segmentation masks for attributes and objects separately, taking the Intersection-over-Union (IoU) of the masks combined with grounding confidence.
    • Spatial Relationship Reward: GLIP grounded open-set object detection.
    • Non-Spatial Relationship Reward: BLIP image captioning followed by CLIP text-text similarity against the prompt.
    • Complex Composition Reward: An aggregate of Grounded-SAM, GLIP, and BLIP-CLIP rewards.
  6. Knowl 6 — GORS Reward-Weighted Diffusion Loss Objective

    equation

    The fine-tuning objective for Generative mOdel finetuning with Reward-driven Sample selection (GORS) scales the standard latent diffusion denoising score-matching loss by the sample alignment reward ss:

    L(θ)=E(x,y,s)∈Ds, t, ϵ∼N(0,I)[s⋅∥ϵ−ϵθ(zt,t,y)∥22]\mathcal{L}(\theta) = \mathbb{E}_{(x, y, s) \in \mathcal{D}_s, \, t, \, \epsilon \sim \mathcal{N}(0, I)} \left[ s \cdot \left\| \epsilon - \epsilon_\theta(z_t, t, y) \right\|_2^2 \right]

    where:

    • Ds={(x,y,s)∣s>τ}\mathcal{D}_s = \{(x, y, s) \mid s > \tau\} is the filtered training dataset of images xx, corresponding prompts yy, and alignment rewards s∈[0,1]s \in [0, 1] exceeding threshold τ\tau.
    • zt=αˉtz0+1−αˉtϵz_t = \sqrt{\bar{\alpha}_t} z_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon is the noisy latent representation of image xx (latent z0z_0) at diffusion timestep t∈{1,…,T}t \in \{1, \dots, T\}.
    • ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is standard Gaussian noise.
    • ϵθ(zt,t,y)\epsilon_\theta(z_t, t, y) is the parameterized noise prediction network conditioned on text prompt yy and timestep tt.
    • ss serves as a per-sample loss weight that prioritizes updates on generated images exhibiting superior semantic compositionality.
  7. Knowl 7 — T2I-CompBench Benchmark Performance Across Text-to-Image Models

    data/table

    Evaluation of six text-to-image models on T2I-CompBench across all six sub-categories. The evaluated models are Stable Diffusion v1-4 (SD v1-4), Stable Diffusion v2 (SD v2), Composable Diffusion v2, Structured Diffusion v2, Attend-and-Excite v2, GORS-unbiased, and GORS. Human evaluation scores are normalized to [0,1][0, 1].

    Category / Sub-category Model CLIP B-CLIP Proposed Metric mGPT-CoT Human
    Color Binding SD v1-4 0.3214 0.7454 0.3765 (B-VQA) 0.7424 0.6533
    SD v2 0.3335 0.7616 0.5065 (B-VQA) 0.7764 0.7747
    Composable v2 0.3178 0.7352 0.4063 (B-VQA) 0.7524 0.6187
    Structured v2 0.3319 0.7626 0.4990 (B-VQA) 0.7822 0.7867
    Attend-and-Excite v2 0.3374 0.7810 0.6400 (B-VQA) 0.8194 0.8240
    GORS-unbiased 0.3390 0.7667 0.6414 (B-VQA) 0.7987 0.8253
    GORS (ours) 0.3395 0.7681 0.6603 (B-VQA) 0.8067 0.8320
    Shape Binding SD v1-4 0.3112 0.7077 0.3576 (B-VQA) 0.7197 0.6160
    SD v2 0.3203 0.7191 0.4221 (B-VQA) 0.7279 0.6587
    Composable v2 0.3092 0.6985 0.3299 (B-VQA) 0.7124 0.5133
    Structured v2 0.3178 0.7177 0.4218 (B-VQA) 0.7228 0.6413
    Attend-and-Excite v2 0.3189 0.7209 0.4517 (B-VQA) 0.7299 0.6360
    GORS-unbiased 0.3175 0.7149 0.4546 (B-VQA) 0.7263 0.6573
    GORS (ours) 0.2973 0.7201 0.4785 (B-VQA) 0.7303 0.7040
    Texture Binding SD v1-4 0.3081 0.7111 0.4156 (B-VQA) 0.7836 0.7227
    SD v2 0.3185 0.7240 0.4922 (B-VQA) 0.7851 0.7827
    Composable v2 0.3092 0.6995 0.3645 (B-VQA) 0.7588 0.6333
    Structured v2 0.3167 0.7234 0.4900 (B-VQA) 0.7806 0.7760
    Attend-and-Excite v2 0.3171 0.7206 0.5963 (B-VQA) 0.8062 0.8400
    GORS-unbiased 0.3216 0.7291 0.6025 (B-VQA) 0.7985 0.8413
    GORS (ours) 0.3233 0.7315 0.6287 (B-VQA) 0.8106 0.8573
    Spatial Relationship SD v1-4 0.3142 0.7667 0.1246 (UniDet) 0.8338 0.3813
    SD v2 0.3206 0.7723 0.1342 (UniDet) 0.8367 0.3467
    Composable v2 0.3001 0.7409 0.0800 (UniDet) 0.8222 0.3080
    Structured v2 0.3201 0.7726 0.1386 (UniDet) 0.8361 0.3467
    Attend-and-Excite v2 0.3213 0.7742 0.1455 (UniDet) 0.8407 0.4027
    GORS-unbiased 0.3237 0.7882 0.1725 (UniDet) 0.8241 0.4467
    GORS (ours) 0.3242 0.7854 0.1815 (UniDet) 0.8362 0.4560
    Non-Spatial Relationship SD v1-4 0.3079 0.7565 — 0.8170 0.9653
    SD v2 0.3127 0.7609 — 0.8235 0.9827
    Composable v2 0.2980 0.7038 — 0.7936 0.8120
    Structured v2 0.3111 0.7614 — 0.8221 0.9773
    Attend-and-Excite v2 0.3109 0.7607 — 0.8214 0.9533
    GORS-unbiased 0.3158 0.7641 — 0.8353 0.9534
    GORS (ours) 0.3193 0.7619 — 0.8172 0.9853
    Complex Compositions SD v1-4 0.2876 0.6816 0.3080 (3-in-1) 0.8075 0.8067
    SD v2 0.3096 0.6893 0.3386 (3-in-1) 0.8094 0.8480
    Composable v2 0.3014 0.6638 0.2898 (3-in-1) 0.8083 0.7520
    Structured v2 0.3084 0.6902 0.3355 (3-in-1) 0.8076 0.8333
    Attend-and-Excite v2 0.2913 0.6875 0.3401 (3-in-1) 0.8078 0.8573
    GORS-unbiased 0.3137 0.6888 0.3470 (3-in-1) 0.8122 0.8654
    GORS (ours) 0.2973 0.6841 0.3328 (3-in-1) 0.8095 0.8680

    Key takeaways:

    1. Spatial relationships represent the most challenging task for text-to-image models (UniDet scores <0.19<0.19, human scores 0.30–0.460.30\text{--}0.46), while non-spatial relationships are the easiest (human scores >0.95>0.95).
    2. GORS achieves top performance across all compositional categories, and GORS-unbiased performs competitively, proving that GORS gains are robust to the choice of sample selection reward model.
  8. Knowl 8 — Human Correlation of Compositional Text-to-Image Evaluation Metrics

    empirical result

    The alignment between automatic evaluation metrics and human perception was evaluated using Kendall's rank correlation (τ\tau) and Spearman's rank correlation (ρ\rho) over 1,800 generated image-prompt pairs on Amazon Mechanical Turk:

    Metric Attr-Color Attr-Shape Attr-Texture Spatial Rel Non-spatial Rel Complex
    τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho
    CLIP 0.1938 0.2773 0.0555 0.0821 0.2890 0.4008 0.2741 0.3548 0.2470 0.3161 0.0650 0.0847
    B-CLIP 0.2674 0.3788 0.1692 0.2413 0.2999 0.4187 0.1983 0.2544 0.2342 0.2964 0.1963 0.2755
    B-VQA-n 0.4602 0.6179 0.2280 0.3180 0.4227 0.5830 — — — — — —
    B-VQA 0.6297 0.7958 0.2707 0.3795 0.5177 0.6995 — — — — — —
    UniDet — — — — — — 0.4756 0.5136 — — — —
    3-in-1 — — — — — — — — — — 0.2831 0.3853
    mGPT 0.1197 0.1616 0.1282 0.1775 0.1061 0.1460 0.0208 0.0229 0.1181 0.1418 0.0066 0.0084
    mGPT-CoT 0.3156 0.4151 0.1300 0.1805 0.3453 0.4664 0.1096 0.1239 0.1944 0.2137 0.1251 0.1463
    mGPT-CLIP 0.2301 0.3174 0.0695 0.0963 0.2004 0.2784 0.1478 0.1950 0.1507 0.1942 0.1457 0.2014

    Key findings:

    1. Disentangled BLIP-VQA (B-VQA) exhibits the highest human correlation on attribute binding, substantially outperforming CLIPScore, BLIP-CLIP, and naive single-question BLIP-VQA (B-VQA-n).
    2. UniDet-based metric correlates best with human assessment for spatial relationships (τ=0.4756,ρ=0.5136\tau=0.4756, \rho=0.5136), whereas holistic vision-language metrics struggle to evaluate spatial layout.
    3. 3-in-1 metric provides the strongest correlation for complex compositions (τ=0.2831,ρ=0.3853\tau=0.2831, \rho=0.3853).
    4. Incorporating Chain-of-Thought into MiniGPT-4 (mGPT-CoT) substantially improves correlation over zero-shot MiniGPT-4 (mGPT) and caption similarity (mGPT-CLIP), though MLLMs still fall short of dedicated specialized metrics.
  9. Knowl 9 — Ablation and Scalability Analysis of GORS Finetuning

    empirical result

    Ablation studies on fine-tuning components, sample selection thresholds, and dataset scaling demonstrate the core design mechanics of GORS:

    1. Finetuning Architecture Targets (Evaluated on Color Binding):

      • U-Net Only with LoRA: Disentangled BLIP-VQA = 0.6216, mGPT-CoT = 0.7840
      • CLIP Text Encoder Only with LoRA: Disentangled BLIP-VQA = 0.5507, mGPT-CoT = 0.7663
      • Both CLIP and U-Net (Full GORS): Disentangled BLIP-VQA = \textbf{0.6570}, mGPT-CoT = \textbf{0.7899} Jointly adapting both the text encoder and the U-Net yields the highest compositionality.
    2. Reward Selection Threshold τ\tau (Evaluated on Color Binding):

      • 0 Threshold (No filtering; fine-tuning on all generated images): Disentangled BLIP-VQA = 0.6130, mGPT-CoT = 0.7879
      • Half Threshold: Disentangled BLIP-VQA = 0.6157, mGPT-CoT = 0.7886
      • Full Threshold (GORS default): Disentangled BLIP-VQA = \textbf{0.6570}, mGPT-CoT = \textbf{0.7899} Selecting only high-reward samples above τ\tau prevents low-quality, misaligned generations from degrading synthesis performance.
    3. Training Dataset Scalability (Evaluated on Complex Compositions using 3-in-1 Metric): Scaling the number of complex composition training prompts from 25 to 1,400 leads to monotonic performance gains:

      • 25 prompts: 0.2596
      • 275 prompts: 0.3086
      • 350 prompts: 0.3299
      • 700 prompts: 0.3328
      • 1,050 prompts: 0.3371
      • 1,400 prompts: \textbf{0.3504}
  10. Knowl 10 — Limitations of T2I-CompBench Metrics and Spatial Evaluation

    limitation

    T2I-CompBench and its evaluation framework have several explicit limitations:

    1. Absence of a Unified Evaluation Metric: No single automatic metric works universally across all compositional dimensions, requiring distinct specialized evaluators (Disentangled BLIP-VQA for attributes, UniDet for spatial layout, CLIP for non-spatial relations, and 3-in-1 for complex scenes).
    2. Failure Cases of Disentangled BLIP-VQA: Disentangled BLIP-VQA fails when objects are partially occluded, when shape or texture descriptions are visually ambiguous or uncommon, or when objects are small and difficult to detect.
    3. 2D Constraint on Spatial Evaluation: The UniDet-based metric evaluates only 2D image-plane bounding box coordinates (x,yx, y centers and IoU) and cannot assess 3D spatial relationships, depth perspective, or complex 3D occlusions.
    4. MLLM Hallucination: Multimodal LLMs (e.g., MiniGPT-4) exhibit visual hallucinations and imperfect cross-modal grounding, limiting their standalone effectiveness as unified compositional evaluators despite Chain-of-Thought prompting.

Coverage note — None was omitted; all primary contributions—including the benchmark prompt design, the specialized evaluation metrics (Disentangled BLIP-VQA, UniDet-based metric, 3-in-1, and MiniGPT4-CoT), the GORS methodology and objective, the extensive benchmarking and human correlation experiments, ablation studies, dataset scaling, and stated limitations—have been included.

References

  1. 1.Robin Rombach et al. “High-resolution image synthesis with latent diffusion models”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 10684–10695.
  2. 2.Jonathan Ho et al. “Cascaded diffusion models for high fidelity image generation”. In: The Journal of Machine Learning Research 23.1 (2022), pp. 2249–2281.
  3. 3.Chitwan Saharia et al. “Photorealistic text-to-image diffusion models with deep language understanding”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 36479–36494.
  4. 4.Prafulla Dhariwal and Alexander Nichol. “Diffusion models beat gans on image synthesis”. In: Advances in neural information processing systems 34 (2021), pp. 8780–8794.
  5. 5.Alexander Quinn Nichol and Prafulla Dhariwal. “Improved denoising diffusion probabilistic models”. In: International Conference on Machine Learning. PMLR. 2021, pp. 8162–8171.
  6. 6.Huiwen Chang et al. “Muse: Text-to-image generation via masked generative transformers”. In: arXiv preprint arXiv:2301.00704 (2023).
  7. 7.Nan Liu et al. “Compositional visual generation with composable diffusion models”. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVII. Springer. 2022, pp. 423–439.
  8. 8.Weixi Feng et al. “Training-Free Structured Diffusion Guidance for Compositional Text-toImage Synthesis”. In: ICLR. 2023.
  9. 9.Hila Chefer et al. “Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models”. In: (2023).
  10. 10.Qiucheng Wu et al. “Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis”. In: arXiv preprint arXiv:2304.03869 (2023).
  11. 11.Alec Radford et al. “Learning transferable visual models from natural language supervision”. In: International conference on machine learning. PMLR. 2021, pp. 8748–8763.
  12. 12.Jack Hessel et al. “Clipscore: A reference-free evaluation metric for image captioning”. In: arXiv preprint arXiv:2104.08718 (2021).
  13. 13.Junnan Li et al. “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation”. In: International Conference on Machine Learning. PMLR. 2022, pp. 12888–12900.
  14. 14.Junnan Li et al. “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models”. In: arXiv preprint arXiv:2301.12597 (2023).
  15. 15.Deyao Zhu et al. “Minigpt-4: Enhancing vision-language understanding with advanced large language models”. In: arXiv preprint arXiv:2304.10592 (2023).
  16. 16.Jason Wei et al. “Chain-of-thought prompting elicits reasoning in large language models”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 24824–24837.
  17. 17.Eslam Mohamed Bakr et al. “HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models”. In: arXiv preprint arXiv:2304.05390 (2023).
  18. 18.Scott Reed et al. “Generative adversarial text to image synthesis”. In: ICML. 2016.
  19. 19.Scott E Reed et al. “Learning what and where to draw”. In: NeurIPS. 2016.
  20. 20.Han Zhang et al. “Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks”. In: ICCV (2017).
  21. 21.Tao Xu et al. “Attngan: Fine-grained text to image generation with attentional generative adversarial networks”. In: CVPR (2018).
  22. 22.Minfeng Zhu et al. “Dm-gan: Dynamic memory generative adversarial networks for text-toimage synthesis”. In: CVPR. 2019.
  23. 23.Han Zhang et al. “Cross-Modal Contrastive Learning for Text-to-Image Generation”. In: CVPR. 2021.
  24. 24.Ian Goodfellow et al. “Generative adversarial nets”. In: 2014.
  25. 25.Aditya Ramesh et al. “Zero-shot text-to-image generation”. In: ICML. 2021.
  26. 26.Aditya Ramesh et al. “Hierarchical text-conditional image generation with clip latents”. In: arXiv preprint arXiv:2204.06125 (2022).
  27. 27.Alex Nichol et al. “Glide: Towards photorealistic image generation and editing with text-guided diffusion models”. In: ICML. 2022.
  28. 28.Chitwan Saharia et al. “Photorealistic text-to-image diffusion models with deep language understanding”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 36479–36494.
  29. 29.Oran Gafni et al. “Make-a-scene: Scene-based text-to-image generation with human priors”. In: ECCV. 2022.
  30. 30.Shu Zhang et al. “HIVE: Harnessing Human Feedback for Instructional Visual Editing”. In: arXiv preprint arXiv:2303.09618 (2023).
  31. 31.Kimin Lee et al. “Aligning text-to-image models using human feedback”. In: arXiv preprint arXiv:2302.12192 (2023).
  32. 32.Hanze Dong et al. “Raft: Reward ranked finetuning for generative foundation model alignment”. In: arXiv preprint arXiv:2304.06767 (2023).
  33. 33.Zhiheng Li et al. “Stylet2i: Toward compositional and high-fidelity text-to-image synthesis”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 18197–18207.
  34. 34.Dong Huk Park et al. “Benchmark for compositional text-to-image synthesis”. In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1). 2021.
  35. 35.Minghao Chen, Iro Laina, and Andrea Vedaldi. “Training-free layout control with crossattention guidance”. In: arXiv preprint arXiv:2304.03373 (2023).
  36. 36.Catherine Wah et al. “The caltech-ucsd birds-200-2011 dataset”. In: (2011).
  37. 37.Maria-Elena Nilsback and Andrew Zisserman. “Automated flower classification over a large number of classes”. In: 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing. IEEE. 2008, pp. 722–729.
  38. 38.Tsung-Yi Lin et al. “Microsoft coco: Common objects in context”. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. Springer. 2014, pp. 740–755.
  39. 39.Jaemin Cho, Abhay Zala, and Mohit Bansal. “Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers”. In: arXiv preprint arXiv:2202.04053 (2022).
  40. 40.Vitali Petsiuk et al. “Human evaluation of text-to-image models on a multi-task benchmark”. In: arXiv preprint arXiv:2211.12112 (2022).
  41. 41.Tim Salimans et al. “Improved techniques for training gans”. In: Advances in neural information processing systems 29 (2016).
  42. 42.Martin Heusel et al. “Gans trained by a two time-scale update rule converge to a local nash equilibrium”. In: Advances in neural information processing systems 30 (2017).
  43. 43.Yujie Lu et al. “LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation”. In: arXiv preprint arXiv:2305.11116 (2023).
  44. 44.Yixiong Chen. “X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models”. In: arXiv preprint arXiv:2305.10843 (2023).
  45. 45.OpenAI. https://openai.com/blog/chatgpt/. 2023.
  46. 46.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. “Simple multi-dataset detection”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 7571–7580.
  47. 47.Edward J Hu et al. “Lora: Low-rank adaptation of large language models”. In: arXiv preprint arXiv:2106.09685 (2021).
  48. 48.Shuai Shao et al. “Objects365: A large-scale, high-quality dataset for object detection”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2019, pp. 8430–8439.
  49. 49.Alina Kuznetsova et al. “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale”. In: International Journal of Computer Vision 128.7 (2020), pp. 1956–1981.
  50. 50.Gerhard Neuhold et al. “The mapillary vistas dataset for semantic understanding of street scenes”. In: Proceedings of the IEEE international conference on computer vision. 2017, pp. 4990–4999.
  51. 51.https://github.com/huggingface/diffusers/blob/main/examples/text_to_image.
  52. 52.Ilya Loshchilov and Frank Hutter. “Decoupled weight decay regularization”. In: arXiv preprint arXiv:1711.05101 (2017).
  53. 53.Shilong Liu et al. “Grounding DINO: Marrying DINO with Grounded Pre-Training for OpenSet Object Detection”. In: arXiv preprint arXiv:2303.05499 (2023).
  54. 54.Liunian Harold Li et al. “Grounded language-image pre-training”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 10965–10975.

Citation

MLA
Huang, K., et al. “T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation”. Advances in Neural Information Processing Systems 36, 2023, pp. 78723–47, https://doi.org/10.52202/075280-3443.
APA
Huang, K., Sun, K., Xie, E., Li, Z., & Liu, X. (2023). T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation. Advances in Neural Information Processing Systems 36, 78723–78747. https://doi.org/10.52202/075280-3443
Chicago
Huang, K., K. Sun, E. Xie, Z. Li, and X. Liu. 2023. “T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation”. Advances in Neural Information Processing Systems 36, 78723–47. https://doi.org/10.52202/075280-3443.
Harvard
Huang, K. et al. (2023) “T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation”, Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 78723–78747. Available at: https://doi.org/10.52202/075280-3443.
Vancouver
1. Huang K, Sun K, Xie E, Li Z, Liu X (2023) T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation. In: Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 78723–78747

BibTeX

@inproceedings{Huang_2023, series={NeurIPS 2023}, title={T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation}, url={http://dx.doi.org/10.52202/075280-3443}, DOI={10.52202/075280-3443}, booktitle={Advances in Neural Information Processing Systems 36}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Huang, Kaiyi and Sun, Kaiyue and Xie, Enze and Li, Zhenguo and Liu, Xihui}, year={2023}, pages={78723–78747}, collection={NeurIPS 2023} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors