VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Max KuDongfu JiangCong WeiXiang YueWenhu Chen

article2024ACL229 citations

Introduces VIEScore, a training-free evaluation metric powered by multimodal large language models that generates human-aligned scores and explanatory rationales across diverse conditional image synthesis tasks.

Listen

Evaluating artificial intelligence systems that generate and edit images has become a critical bottleneck in visual computing. Traditional automated metrics produce single opaque numbers that fail to explain why an image passed or failed, and they cannot adapt flexibly across different generation and editing tasks. While human evaluations provide nuanced judgment, they are costly, difficult to scale, and subjective. To address these challenges, the article evaluates VIEScore, a visual instruction-guided evaluation framework that leverages multimodal large language models to assess conditional image synthesis without requiring extra training or fine-tuning.

The main objective of the article is to demonstrate whether advanced multimodal language models can serve as task-aware, explainable automated evaluators that correlate closely with human visual assessments across diverse image generation and editing scenarios. The researchers evaluated the framework across seven core image synthesis tasks spanning 29 generative models and more than 14,000 human ratings. The evaluation prompts separate visual assessment into explicit sub-scores for semantic consistency (alignment with conditions) and perceptual quality (naturalness and absence of artifacts), requiring the underlying model to generate natural-language rationales before outputting numeric scores.

The analysis revealed several critical findings. First, advanced proprietary models demonstrated strong alignment with human judgments: VIEScore powered by GPT-4o achieved an overall Spearman correlation of 0.40 with human ratings across all tasks, approaching the human-to-human agreement baseline of 0.46. Second, current open-source multimodal models fell significantly behind; models such as LLaVA, BLIP-2, and Qwen-VL frequently failed to follow complex instructions or generated uninformative score distributions, yielding overall correlations below 0.10. Third, multimodal evaluators performed substantially better on image generation tasks than on image editing tasks, largely because they frequently missed subtle, fine-grained modifications like localized texture or color shifts. Finally, providing example images via in-context learning consistently degraded evaluation performance by confusing the models, whereas removing condition inputs during perceptual quality checks improved human correlation.

These findings suggest that state-of-the-art multimodal models can effectively replace costly human evaluations for standard image generation tasks, significantly reducing evaluation overhead and accelerating model benchmarking. The natural-language explanations generated by the system also improve trust and transparency in automated quality assurance. However, practitioners must remain cautious when deploying automated evaluators for image editing tasks or multi-image reasoning contexts, where subtle artifacts may go undetected.

Organizations evaluating generative visual models should adopt zero-shot multimodal prompting frameworks rather than few-shot image examples, isolating visual quality checks from prompt inputs to maximize accuracy. Future development should focus on building lightweight, distilled evaluator models capable of replicating human-level inspection at a lower cost, alongside improving multimodal sensitivity to fine-grained visual edits. Leaders should note that these findings rely on current proprietary model APIs, which are subject to content-filtering restrictions that omit photorealistic human subjects and introduce platform dependencies.

Cover for VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Abstract

In the rapidly advancing field of conditional image generation research, challenges such as limited explainability lie in effectively evaluating the performance and capabilities of various models. This paper introduces VIESCORE, a Visual Instruction-guided Explainable metric for evaluating any conditional image generation tasks. VIESCORE leverages general knowledge from Multimodal Large Language Models (MLLMs) as the backbone and does not require training or fine-tuning. We evaluate VIESCORE on seven prominent tasks in conditional image tasks and found: (1) VIESCORE (GPT4-o) achieves a high Spearman correlation of 0.4 with human evaluations, while the human-to-human correlation is 0.45. (2) VIESCORE (with open-source MLLM) is significantly weaker than GPT-4o and GPT-4v in evaluating synthetic images. (3) VIESCORE achieves a correlation on par with human ratings in the generation tasks but struggles in editing tasks. With these results, we believe VIESCORE shows its great potential to replace human judges in evaluating image synthesis tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Conditional Image Synthesis
  • 2.2 Synthetic Images Evaluation
  • 2.3 Large Language Models as Evaluators
  • 3 Preliminary
  • 3.1 Evaluation Benchmark
  • 3.2 Multimodal Large Language Models
  • 3.3 Existing Auto-metrics
  • 4 Method
  • 5 Experimental Results
  • 5.1 Correlation Study
  • 5.2 Insights and Challenges on VIESCORE
  • 5.3 VIESCORE and Auto-metrics vs Human
  • 6 Conclusion
  • 7 Limitations
  • 8 Potential Risks
  • 9 Artifacts
  • 10 Computational Experiments
  • 11 Acknowledgement
  • References
  • A Prompt Templates
  • B Supplementary Information
  • B.1 Human Correlation Study
  • B.2 Zero-shot vs One-shot on VIE
  • B.3 Autometrics vs Human Detail results
  • B.4 ImagenHub Models used
  • B.5 ImagenHub Human data information
  • C Backbone Performances
  • C.1 Parsing MLLM outputs
  • C.2 Observations of GPT-4o.
  • C.3 Observations of Gemini-Pro.
  • C.4 Observations of GPT-4v.
  • C.5 Observations of LLaVA.
  • C.6 Observations of Qwen-VL.
  • C.7 Observations of BLIP2.
  • C.8 Observations of InstructBLIP.
  • C.9 Observations of Fuyu.
  • C.10 Observations of CogVLM.
  • C.11 Observations of OpenFlamingo.

Knowls

  1. Knowl 1 — VIEScore Evaluation Framework and Scoring Formulation

    model/method

    VIEScore is a visual instruction-guided, training-free evaluation framework that uses Multimodal Large Language Models (MLLMs) to assess conditional image synthesis tasks by generating natural language rationales alongside numerical scores. Given an evaluation instruction II, a synthesized image OO, and a set of input conditions C∗C^* (which can include text prompts, reference concept images, or control guidance maps), the evaluation function is defined as:

    fVIE(I,O,C∗)=(rationale,score)f_{\text{VIE}}(I, O, C^*) = (\text{rationale}, \text{score})

    The framework separates image assessment into two complementary rating dimensions:

    1. Semantic Consistency (SC): Evaluates how faithfully the synthesized image aligns with all input conditions C∗C^*. Both C∗C^* and OO are fed simultaneously into the MLLM. SC is decomposed into task-specific sub-scores α1,…,αn∈[0,10]\alpha_1, \dots, \alpha_n \in [0, 10].
    2. Perceptual Quality (PQ): Evaluates image realism, artifacts, and naturalness. Only the synthesized image OO is fed into the MLLM (excluding C∗C^*) to prevent the model from conflating condition alignment with visual realism. PQ is decomposed into sub-scores β1,…,βm∈[0,10]\beta_1, \dots, \beta_m \in [0, 10] measuring naturalness and artifact presence.

    To ensure that a synthesized image failing on any individual requirement is penalised, the overall instance-level score OoverallO_{\text{overall}} is computed using the geometric mean of the minimum sub-scores across each dimension:

    Ooverall=[min⁡(α1,…,αn)⋅min⁡(β1,…,βm)]1/2O_{\text{overall}} = \left[ \min(\alpha_1, \dots, \alpha_n) \cdot \min(\beta_1, \dots, \beta_m) \right]^{1/2}

    Sub-scores on the 0 to 10 scale are normalized to the range [0.0,1.0][0.0, 1.0] when comparing with human ratings.

  2. Knowl 2 — Task-Decomposed Instruction Prompting in VIEScore

    model/method

    VIEScore prompts MLLMs using a two-segment template consisting of a general context prompt and a task-specific rating prompt. To ensure consistent machine parsing across MLLMs, the model is prompted to return a JSON object with structured keys:

    {
      "score": [...],
      "reasoning": "..."
    }
    

    The scoring instructions decompose high-level criteria into distinct integer sub-scores on a 0 to 10 scale:

    • Perceptual Quality (PQ) [all tasks]:
      • Naturalness (β1\beta_1): 0 denotes completely unnatural scene layout, improper lighting, distance perspective, or inconsistent shadows; 10 denotes a completely natural scene.
      • Artifacts (β2\beta_2): 0 denotes severe distortion, watermarks, scratches, blurred facial features, or unharmonized subjects; 10 denotes an image completely free of artifacts. Output format: [naturalness, artifacts].
    • Semantic Consistency (SC):
      • Text-guided Image Generation: Single score α1∈[0,10]\alpha_1 \in [0, 10] evaluating prompt adherence.
      • Text/Mask-guided Image Editing: Two sub-scores: editing execution success α1∈[0,10]\alpha_1 \in [0, 10] and degree of overediting α2∈[0,10]\alpha_2 \in [0, 10] (where 10 indicates minimal unnecessary change from the source image).
      • Control-guided Image Generation: Two sub-scores: text prompt alignment α1∈[0,10]\alpha_1 \in [0, 10] and adherence to the structural guidance map α2∈[0,10]\alpha_2 \in [0, 10] (e.g., Canny edge, OpenPose).
      • Subject-driven Image Generation: Two sub-scores: prompt alignment α1∈[0,10]\alpha_1 \in [0, 10] and visual resemblance to the reference token subject image α2∈[0,10]\alpha_2 \in [0, 10].
      • Subject-guided Image Editing: Two sub-scores: token subject resemblance α1∈[0,10]\alpha_1 \in [0, 10] and preservation/overediting control relative to the source image α2∈[0,10]\alpha_2 \in [0, 10].
  3. Knowl 3 — Benchmark Correlation of MLLMs on Conditional Image Synthesis Evaluation

    empirical result

    Across 7 conditional image synthesis tasks evaluated on 29 generative models using 14,403 human ratings from the ImagenHub benchmark, MLLM backbones exhibit marked variations in Spearman correlation with human raters (averaged across models and tasks using Fisher Z-transformation):

    Backbone / Evaluator M-HcorrSCM\text{-}H^{\text{SC}}_{\text{corr}} M-HcorrPQM\text{-}H^{\text{PQ}}_{\text{corr}} M-HcorrOM\text{-}H^{\text{O}}_{\text{corr}}
    Human Raters (Inter-Rater Agreement) 0.4700 0.4124 0.4558
    VIEScore (GPT-4o, 0-shot) 0.4459 0.3399 0.4041
    VIEScore (GPT-4o, 1-shot) 0.4309 0.1167 0.3770
    VIEScore (Gemini-1.5-Pro, 0-shot) 0.3322 0.2675 0.3048
    VIEScore (Gemini-1.5-Pro, 1-shot) 0.3094 0.3070 0.3005
    VIEScore (GPT-4v, 0-shot) 0.3655 0.3092 0.3266
    VIEScore (GPT-4v, 1-shot) 0.2689 0.2338 0.2604
    VIEScore (LLaVA-1.5-7B, 0-shot) 0.1046 0.0319 0.0925
    VIEScore (LLaVA-1.5-7B, 1-shot) 0.1012 0.0138 0.0695
    VIEScore (Qwen-VL-7B, 0-shot) 0.0679 0.0165 0.0920
    VIEScore (BLIP-2, 0-shot) 0.0504 -0.0108 0.0622
    VIEScore (InstructBLIP, 0-shot) 0.0246 0.0095 0.0005
    VIEScore (Fuyu-8B, 0-shot) -0.0110 -0.0172 0.0154
    VIEScore (CogVLM, 0-shot) -0.0228 0.0514 -0.0050
    VIEScore (OpenFlamingo, 0-shot) -0.0037 -0.0102 -0.0122

    where M-HcorrSCM\text{-}H^{\text{SC}}_{\text{corr}}, M-HcorrPQM\text{-}H^{\text{PQ}}_{\text{corr}}, and M-HcorrOM\text{-}H^{\text{O}}_{\text{corr}} denote Metric-to-Human Spearman correlations for Semantic Consistency, Perceptual Quality, and Overall score, respectively.

    GPT-4o achieves the highest correlation (0.40410.4041 overall), nearing human-to-human agreement (0.45580.4558). Closed-source proprietary models (GPT-4o, GPT-4v, Gemini-1.5-Pro) substantially outperform open-source models, which fail to adhere to format instructions or produce degenerate score distributions.

  4. Knowl 4 — Performance Degradation from In-Context Few-Shot Examples in Visual Evaluation

    empirical result

    Applying In-Context Learning (ICL) by providing one-shot demonstration examples (an exemplar image pair, reference scores, and a written rationale) causes an overall decline in evaluation correlation across MLLMs rather than an improvement:

    • GPT-4o overall Metric-to-Human Spearman correlation across all 7 tasks drops from 0.40410.4041 (0-shot) to 0.37700.3770 (1-shot), with Perceptual Quality correlation plummeting from 0.33990.3399 to 0.11670.1167.
    • GPT-4v overall correlation drops from 0.32660.3266 (0-shot) to 0.26040.2604 (1-shot), with Semantic Consistency dropping from 0.36550.3655 to 0.26890.2689.
    • In Multi-Concept Image Composition, GPT-4o correlation falls from 0.41360.4136 to 0.35230.3523, and GPT-4v falls from 0.33460.3346 to 0.19180.1918.
    • In Subject-Driven Image Editing, GPT-4o correlation falls from 0.32680.3268 to 0.27970.2797, and GPT-4v falls from 0.15070.1507 to −0.0139-0.0139.

    Qualitative inspection of generated rationales demonstrates that MLLMs become confused by the visual features of the in-context demonstration images, inadvertently evaluating elements from the example image rather than restricting reasoning to the query image.

  5. Knowl 5 — Isolation of Input Conditions in Perceptual Quality Assessment

    empirical result

    Feeding input conditions (such as input source images or reference concept images) alongside the synthesized image into the MLLM prompt during Perceptual Quality (PQ) evaluation significantly lowers correlation with human evaluations:

    PQ Prompting Setting (GPT-4v) Text-guided Image Editing (M-HcorrPQM\text{-}H^{\text{PQ}}_{\text{corr}}) Multi-concept Image Composition (M-HcorrPQM\text{-}H^{\text{PQ}}_{\text{corr}})
    Human Raters (with inputs) 0.5052 0.5145
    Without inputs (synthetic image only) 0.4274 0.3025
    With inputs (conditions included) 0.2256 0.0731

    Removing input conditions from the prompt isolates the evaluation strictly to visual naturalness, rendering quality, and absence of visual artifacts, preventing MLLMs from conflating conditional semantic fidelity with intrinsic perceptual quality.

  6. Knowl 6 — Comparison of VIEScore with Traditional Evaluation Metrics Across Tasks

    empirical result

    Across seven conditional image generation and editing tasks, VIEScore demonstrates stronger alignment with human ratings than task-agnostic automated metrics (CLIP-Score, LPIPS, DINO, CLIP-I):

    • Text-guided Image Generation: CLIP-Score correlates negatively with human ratings (M-HcorrSC=−0.0817M\text{-}H^{\text{SC}}_{\text{corr}} = -0.0817, M-HcorrO=−0.0881M\text{-}H^{\text{O}}_{\text{corr}} = -0.0881), clustering in a narrow score range ([0.25,0.35][0.25, 0.35]) that fails to distinguish quality variations. GPT-4o achieves M-HcorrSC=0.4989M\text{-}H^{\text{SC}}_{\text{corr}} = 0.4989 and M-HcorrO=0.3928M\text{-}H^{\text{O}}_{\text{corr}} = 0.3928 (human inter-rater correlation is 0.46520.4652).
    • Subject-driven Image Generation and Editing: DINO embeddings align well with single-subject identity (M-HcorrSC=0.4160M\text{-}H^{\text{SC}}_{\text{corr}} = 0.4160, M-HcorrO=0.4246M\text{-}H^{\text{O}}_{\text{corr}} = 0.4246 on generation), outperforming GPT-4v (0.37380.3738) and CLIP-I (0.30580.3058). However, GPT-4o achieves the highest human correlation overall (M-HcorrSC=0.4806M\text{-}H^{\text{SC}}_{\text{corr}} = 0.4806, M-HcorrO=0.4637M\text{-}H^{\text{O}}_{\text{corr}} = 0.4637).
    • Control-guided Image Generation: LPIPS effectively captures structural distortions (M-HcorrO=0.4133M\text{-}H^{\text{O}}_{\text{corr}} = 0.4133), outperforming GPT-4v (0.39990.3999) and Gemini-Pro (0.29600.2960), but GPT-4o outperforms all baselines with M-HcorrO=0.5439M\text{-}H^{\text{O}}_{\text{corr}} = 0.5439 (exceeding human agreement of 0.53070.5307).
    • Image Editing Tasks: LPIPS performs poorly on mask-guided editing (M-HcorrO=−0.0694M\text{-}H^{\text{O}}_{\text{corr}} = -0.0694) and text-guided editing (M-HcorrO=0.1142M\text{-}H^{\text{O}}_{\text{corr}} = 0.1142) because modern editing models produce artifact-free outputs where distortion metrics cannot evaluate semantic naturalness or overediting. In contrast, GPT-4o achieves M-HcorrO=0.4769M\text{-}H^{\text{O}}_{\text{corr}} = 0.4769 on mask-guided editing and 0.38210.3821 on text-guided editing.
  7. Knowl 7 — Model-Level Ranking Alignment Between MLLMs and Human Benchmarks

    empirical result

    When aggregating instance-level evaluation scores to rank generative models on the ImagenHub leaderboard, MLLM-based evaluations correlate positively with human rankings across all 7 evaluation categories:

    Spearman's footrule dSF(rHuman,rMethod)↓d_{\text{SF}}(r_{\text{Human}}, r_{\text{Method}})\downarrow Spearman's rho ρS(rHuman,rMethod)↑\rho_{\text{S}}(r_{\text{Human}}, r_{\text{Method}})\uparrow
    Task (Model Count) GPT-4v LLaVA LPIPS CLIP GPT-4v LLaVA LPIPS CLIP
    Text-guided Image Generation (5) 2 6 N/A 8 0.90 0.50 N/A -0.20
    Mask-guided Image Editing (4) 2 8 2 0 0.80 -1.00 0.80 1.00
    Text-guided Image Editing (8) 12 16 20 16 0.67 0.48 0.17 0.48
    Subject-driven Image Generation (4) 4 6 0 6 0.20 -0.40 1.00 -0.20
    Subject-driven Image Editing (3) 2 2 4 4 0.50 0.50 -0.50 -1.00
    Multi-concept Image Composition (3) 0 0 2 2 1.00 1.00 0.50 0.50
    Control-guided Image Generation (2) 0 0 0 2 1.00 1.00 1.00 -1.00

    where dSF(r,r∗)∈[0,+∞)d_{\text{SF}}(r, r^*) \in [0, +\infty) is Spearman's footrule distance and ρS(r,r∗)∈[−1,1]\rho_{\text{S}}(r, r^*) \in [-1, 1] is Spearman's rank correlation coefficient.

    GPT-4v achieves positive rank correlation (ρS≥0.20\rho_{\text{S}} \ge 0.20) across all 7 tasks, achieving exact alignment (ρS=1.00\rho_{\text{S}} = 1.00, dSF=0d_{\text{SF}} = 0) on Multi-concept Image Composition and Control-guided Image Generation. In comparison, CLIP-Score produces negative model-ranking correlations on three of the tasks.

  8. Knowl 8 — MLLM Sensitivity Limitations on Fine-Grained Nuances in Image Editing

    limitation

    MLLMs (including GPT-4o, GPT-4v, Gemini-1.5-Pro, and LLaVA) exhibit lower correlation with human raters on image editing tasks compared to image generation tasks. In text-guided and subject-driven image editing, MLLMs frequently fail to detect localized, fine-grained modifications—such as small patch edits, subtle texture changes, or color alterations. Consequently, MLLMs often judge an edited image as identical to its source image even when human annotators evaluate the edit as successful. This deficiency arises because standard MLLM visual backbones prioritize global high-level semantics over fine-grained spatial and pixel-level details.

  9. Knowl 9 — Safety Filter Refusals in Automated MLLM Image Evaluation

    limitation

    When employing closed-source MLLM APIs (such as GPT-4v) as automated evaluators, strict security and privacy policies cause the model to refuse evaluation on synthetic images that resemble real individuals or human portraits. The model returns standardized refusal responses (e.g., "I am sorry, but I cannot process these images as they contain real people."). In automated benchmarking pipelines, these samples must be detected and dropped via string matching, introducing dataset attrition and potential evaluation bias when assessing models on photorealistic portrait generation.

Coverage note — Detailed per-model breakdown tables for individual baseline metrics (Tables 7, 8, 9, 10, 11, 12, 13) were synthesized into the comparative knowls rather than extracted as repetitive separate tables.

References

  1. 1.Omri Avrahami, Dani Lischinski, and Ohad Fried. 2022. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18208–18218.
  2. 2.Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023. Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv preprint arXiv:2308.01390.
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966.
  4. 4.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. 2022. Text2live: Text-driven layered image and video editing. In European Conference on Computer Vision, pages 707–723.
  5. 5.Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sagnak Ta¸sırlar. 2023. ˘ Introducing our multimodal models.
  6. 6.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. 2023. Instructpix2pix: Learning to follow image editing instructions. In CVPR.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  8. 8.Mathilde Caron, Hugo Touvron, Ishan Misra, Herv’e J’egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640.
  9. 9.Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Rui, Xuhui Jia, Ming-Wei Chang, and William W Cohen. 2023. Subject-driven text-to-image generation via apprenticeship learning. arXiv preprint arXiv:2304.00186.
  10. 10.Jaemin Cho, Yushi Hu, Roopal Garg, Peter Anderson, Ranjay Krishna, Jason Baldridge, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. 2023. Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-image generation. arXiv preprint arXiv:2310.18235.
  11. 11.Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. 2022. Diffedit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations.
  12. 12.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. ArXiv, abs/2305.06500.
  13. 13.deep floyd.ai. 2023. If by deepfloyd lab at stabilityai.
  14. 14.Emily L Denton, Soumith Chintala, Rob Fergus, et al. 2015. Deep generative image models using a laplacian pyramid of adversarial networks. Advances in neural information processing systems, 28.
  15. 15.Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. In Advances in Neural Information Processing Systems, volume 34, pages 8780–8794. Curran Associates, Inc.
  16. 16.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback.
  17. 17.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023a. Gptscore: Evaluate as you desire. ArXiv, abs/2302.04166.
  18. 18.Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. 2023b. Dreamsim: Learning new dimensions of human visual similarity using synthetic data.
  19. 19.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. In The Eleventh International Conference on Learning Representations.
  20. 20.Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. 2016. Deep learning, volume 1. MIT Press.
  21. 21.Jing Gu, Yilin Wang, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong, Qing Liu, Zhifei Zhang, He Zhang, Jianming Zhang, HyunJoon Jung, et al. 2023. Photoswap: Personalized subject swapping in images. arXiv preprint arXiv:2305.18286.
  22. 22.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. Clipscore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7514–7528.
  23. 23.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30.
  24. 24.Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pages 6840–6851. Curran Associates, Inc.
  25. 25.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  26. 26.Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. 2023. Tifa: Accurate and interpretable text-toimage faithfulness evaluation with question answering. arXiv preprint arXiv:2303.11897.
  27. 27.Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. 2023. T2i-compbench: A comprehensive benchmark for open-world compositional textto-image generation. ArXiv, abs/2307.06350.
  28. 28.Mert Inan, Piyush Sharma, Baber Khalid, Radu Soricut, Matthew Stone, and Malihe Alikhani. 2021. COSMic: A coherence-aware generation metric for image descriptions. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3419–3430, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  29. 29.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134.
  30. 30.Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023. Tigerscore: Towards building explainable metric for all text generation tasks. arXiv preprint arXiv:2310.00752.
  31. 31.Jin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo, and Sang-Woo Lee. 2022. Mutual information divergence: A unified metric for multimodal generative models. Advances in Neural Information Processing Systems, 35:35072–35086.
  32. 32.Younghyun Kim, Sangwoo Mo, Min-Kyung Kim, Kyungmin Lee, Jaeho Lee, and Jinwoo Shin. 2023. Bias-to-text: Debiasing unknown visual biases through language interpretation. ArXiv.
  33. 33.Max Ku, Tianle Li, Kai Zhang, Yujie Lu, Xingyu Fu, Wenwen Zhuang, and Wenhu Chen. 2023. Imagenhub: Standardizing the evaluation of conditional image generation models. arXiv preprint arXiv:2310.01596.
  34. 34.Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. 2023. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1931–1941.
  35. 35.Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2019. Improved precision and recall metric for assessing generative models. Advances in Neural Information Processing Systems, 32.
  36. 36.Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, and Kyomin Jung. 2020. ViLBERTScore: Evaluating image caption using vision-and-language BERT. In Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems, pages 34–39, Online. Association for Computational Linguistics.
  37. 37.Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, et al. 2023. Holistic evaluation of textto-image models. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  38. 38.Dongxu Li, Junnan Li, and Steven CH Hoi. 2023a. Blipdiffusion: Pre-trained subject representation for controllable text-to-image generation and editing. arXiv preprint arXiv:2305.14720.
  39. 39.Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023b. Generative judge for evaluating alignment. arXiv preprint arXiv:2310.05470.
  40. 40.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pretraining for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR.
  41. 41.Tianle Li, Max Ku, Cong Wei, and Wenhu Chen. 2023c. Dreamedit: Subject-driven image editing. arXiv preprint arXiv:2306.12624.
  42. 42.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a. Visual instruction tuning. ArXiv, abs/2304.08485.
  43. 43.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? In Workshop on Knowledge Extraction and Integration for Deep Learning Architectures; Deep Learning Inside Out.
  44. 44.Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. 2023b. Cones 2: Customizable image synthesis with multiple subjects. arXiv preprint arXiv:2305.19327.
  45. 45.Lingxiao Lu, Bo Zhang, and Li Niu. 2023a. Dreamcom: Finetuning text-guided inpainting model for image composition. ArXiv, abs/2309.15508.
  46. 46.Yujie Lu, Xiujun Li, William Yang Wang, and Yejin Choi. 2023b. Vim: Probing multimodal large language models for visual embedded instruction following. ArXiv, abs/2311.17647.
  47. 47.Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. 2023c. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation. arXiv preprint arXiv:2305.11116.
  48. 48.Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. 2023d. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation. ArXiv, abs/2305.11116.
  49. 49.Andreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion probabilistic models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11451–11461.
  50. 50.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2021. Sdedit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations.
  51. 51.Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2023. Null-text inversion for editing real images using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6038–6047.
  52. 52.OpenAI. 2023. Gpt-4 technical report.
  53. 53.openjourney.ai. 2023. Openjourney is an open source stable diffusion fine tuned model on midjourney images.
  54. 54.Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. 2021. Benchmark for compositional text-to-image synthesis. In Thirtyfifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  55. 55.Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu. 2023. Zero-shot image-to-image translation. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–11.
  56. 56.Archit Parnami and Minwoo Lee. 2022. Learning from few examples: A summary of approaches to few-shot learning. ArXiv, abs/2203.04291.
  57. 57.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. ArXiv, abs/2306.14824.
  58. 58.Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, et al. 2023. Unicontrol: A unified diffusion model for controllable visual generation in the wild. arXiv preprint arXiv:2305.11147.
  59. 59.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical textconditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3.
  60. 60.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. Highresolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695.
  61. 61.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510.
  62. 62.runwayml. 2023. Stable diffusion inpainting.
  63. 63.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494.
  64. 64.Mehdi SM Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. 2018. Assessing generative models via precision and recall. Advances in neural information processing systems, 31.
  65. 65.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans. Advances in neural information processing systems, 29.
  66. 66.Ayush Sarkar, Hanlin Mai, Amitabh Mahapatra, Svetlana Lazebnik, David Forsyth, and Anand Bhattad. 2023. Shadows don’t lie and lines can’t bend! generative models don’t know projective geometry...for now. ArXiv, abs/2311.17138.
  67. 67.Shelly Sheynin, Adam Polyak, Uriel Singer, Yuval Kirstain, Amit Zohar, Oron Ashual, Devi Parikh, and Yaniv Taigman. 2023. Emu edit: Precise image editing via recognition and generation tasks. arXiv preprint arXiv:2311.10089.
  68. 68.stability.ai. 2023. Stable diffusion xl.
  69. 69.Gemini Team, Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry, Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew Dai, Katie Millican, Ethan Dyer, Mia Glaese, Thibault Sottiaux, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, James Molloy, Jilin Chen, Michael Isard, Paul Barham, Tom Hennigan, Ross McIlroy, Melvin Johnson, Johan Schalkwyk, Eli Collins, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, Clemens Meyer, Gregory Thornton, Zhen Yang, Henryk Michalewski, Zaheer Abbas, Nathan Schucher, Ankesh Anand, Richard Ives, James Keeling, Karel Lenc, Salem Haykal, Siamak Shakeri, Pranav Shyam, Aakanksha Chowdhery, Roman Ring, Stephen Spencer, Eren Sezener, Luke Vilnis, Oscar Chang, Nobuyuki Morioka, George Tucker, Ce Zheng, Oliver Woodman, Nithya Attaluri, Tomas Kocisky, Evgenii Eltyshev, Xi Chen, Timothy Chung, Vittorio Selo, Siddhartha Brahma, Petko Georgiev, Ambrose Slone, Zhenkai Zhu, James Lottes, Siyuan Qiao, Ben Caine, Sebastian Riedel, Alex Tomala, Martin Chadwick, Juliette Love, Peter Choy, Sid Mittal, Neil Houlsby, Yunhao Tang, Matthew Lamm, Libin Bai, Qiao Zhang, Luheng He, Yong Cheng, Peter Humphreys, Yujia Li, Sergey Brin, Albin Cassirer, Yingjie Miao, Lukas Zilka, Taylor Tobin, Kelvin Xu, Lev Proleev, Daniel Sohn, Alberto Magni, Lisa Anne Hendricks, Isabel Gao, Santiago Ontanon, Oskar Bunyan, Nathan Byrd, Abhanshu Sharma, Biao Zhang, Mario Pinto, Rishika Sinha, Harsh Mehta, Dawei Jia, Sergi Caelles, Albert Webson, Alex Morris, Becca Roelofs, Yifan Ding, Robin Strudel, Xuehan Xiong, Marvin Ritter, Mostafa Dehghani, Rahma Chaabouni, Abhijit Karmarkar, Guangda Lai, Fabian Mentzer, Bibo Xu, YaGuang Li, Yujing Zhang, Tom Le Paine, Alex Goldin, Behnam Neyshabur, Kate Baumli, Anselm Levskaya, Michael Laskin, Wenhao Jia, Jack W. Rae, Kefan Xiao, Antoine He, Skye Giordano, Lakshman Yagati, Jean-Baptiste Lespiau, Paul Natsev, Sanjay Ganapathy, Fangyu Liu, Danilo Martins, Nanxin Chen, Yunhan Xu, Megan Barnes, Rhys May, Arpi Vezer, Junhyuk Oh, Ken Franko, Sophie Bridgers, Ruizhe Zhao, Boxi Wu, Basil Mustafa, Sean Sechrist, Emilio Parisotto, Thanumalayan Sankaranarayana Pillai, Chris Larkin, Chenjie Gu, Christina Sorokin, Maxim Krikun, Alexey Guseynov, Jessica Landon, Romina Datta, Alexander Pritzel, Phoebe Thacker, Fan Yang, Kevin Hui, Anja Hauth, Chih-Kuan Yeh, David Barker, Justin Mao-Jones, Sophia Austin, Hannah Sheahan, Parker Schuh, James Svensson, Rohan Jain, Vinay Ramasesh, Anton Briukhov, DaWoon Chung, Tamara von Glehn, Christina Butterfield, Priya Jhakra, Matthew Wiethoff, Justin Frye, Jordan Grimstad, Beer Changpinyo, Charline Le Lan, Anna Bortsova, Yonghui Wu, Paul Voigtlaender, Tara Sainath, Shane Gu, Charlotte Smith, Will Hawkins, Kris Cao, James Besley, Srivatsan Srinivasan, Mark Omernick, Colin Gaffney, Gabriela Surita, Ryan Burnell, Bogdan Damoc, Junwhan Ahn, Andrew Brock, Mantas Pajarskas, Anastasia Petrushkina, Seb Noury, Lorenzo Blanco, Kevin Swersky, Arun Ahuja, Thi Avrahami, Vedant Misra, Raoul de Liedekerke, Mariko Iinuma, Alex Polozov, Sarah York, George van den Driessche, Paul Michel, Justin Chiu, Rory Blevins, Zach Gleicher, Adrià Recasens, Alban Rrustemi, Elena Gribovskaya, Aurko Roy, Wiktor Gworek, Sébastien M. R. Arnold, Lisa Lee, James Lee-Thorp, Marcello Maggioni, Enrique Piqueras, Kartikeya Badola, Sharad Vikram, Lucas Gonzalez, Anirudh Baddepudi, Evan Senter, Jacob Devlin, James Qin, Michael Azzam, Maja Trebacz, Martin Polacek, Kashyap Krishnakumar, Shuo yiin Chang, Matthew Tung, Ivo Penchev, Rishabh Joshi, Kate Olszewska, Carrie Muir, Mateo Wirth, Ale Jakse Hartman, Josh Newlan, Sheleem Kashem, Vijay Bolina, Elahe Dabir, Joost van Amersfoort, Zafarali Ahmed, James Cobon-Kerr, Aishwarya Kamath, Arnar Mar Hrafnkelsson, Le Hou, Ian Mackinnon, Alexandre Frechette, Eric Noland, Xiance Si, Emanuel Taropa, Dong Li, Phil Crone, Anmol Gulati, Sébastien Cevey, Jonas Adler, Ada Ma, David Silver, Simon Tokumine, Richard Powell, Stephan Lee, Kiran Vodrahalli, Samer Hassan, Diana Mincu, Antoine Yang, Nir Levine, Jenny Brennan, Mingqiu Wang, Sarah Hodkinson, Jeffrey Zhao, Josh Lipschultz, Aedan Pope, Michael B. Chang, Cheng Li, Laurent El Shafey, Michela Paganini, Sholto Douglas, Bernd Bohnet, Fabio Pardo, Seth Odoom, Mihaela Rosca, Cicero Nogueira dos Santos, Kedar Soparkar, Arthur Guez, Tom Hudson, Steven Hansen, Chulayuth Asawaroengchai, Ravi Addanki, Tianhe Yu, Wojciech Stokowiec, Mina Khan, Justin Gilmer, Jaehoon Lee, Carrie Grimes Bostock, Keran Rong, Jonathan Caton, Pedram Pejman, Filip Pavetic, Geoff Brown, Vivek Sharma, Mario Luciˇ c, Rajkumar Samuel, Josip ´ Djolonga, Amol Mandhane, Lars Lowe Sjösund, Elena Buchatskaya, Elspeth White, Natalie Clay, Jiepu Jiang, Hyeontaek Lim, Ross Hemsley, Zeyncep Cankara, Jane Labanowski, Nicola De Cao, David Steiner, Sayed Hadi Hashemi, Jacob Austin, Anita Gergely, Tim Blyth, Joe Stanton, Kaushik Shivakumar, Aditya Siddhant, Anders Andreassen, Carlos Araya, Nikhil Sethi, Rakesh Shivanna, Steven Hand, Ankur Bapna, Ali Khodaei, Antoine Miech, Garrett Tanzer, Andy Swing, Shantanu Thakoor, Lora Aroyo, Zhufeng Pan, Zachary Nado, Jakub Sygnowski, Stephanie Winkler, Dian Yu, Mohammad Saleh, Loren Maggiore, Yamini Bansal, Xavier Garcia, Mehran Kazemi, Piyush Patil, Ishita Dasgupta, Iain Barr, Minh Giang, Thais Kagohara, Ivo Danihelka, Amit Marathe, Vladimir Feinberg, Mohamed Elhawaty, Nimesh Ghelani, Dan Horgan, Helen Miller, Lexi Walker, Richard Tanburn, Mukarram Tariq, Disha Shrivastava, Fei Xia, Qingze Wang, ChungCheng Chiu, Zoe Ashwood, Khuslen Baatarsukh, Sina Samangooei, Raphaël Lopez Kaufman, Fred Alcober, Axel Stjerngren, Paul Komarek, Katerina Tsihlas, Anudhyan Boral, Ramona Comanescu, Jeremy Chen, Ruibo Liu, Chris Welty, Dawn Bloxwich, Charlie Chen, Yanhua Sun, Fangxiaoyu Feng, Matthew Mauger, Xerxes Dotiwalla, Vincent Hellendoorn, Michael Sharman, Ivy Zheng, Krishna Haridasan, Gabe Barth-Maron, Craig Swanson, Dominika Rogozinska, Alek Andreev, Paul Kishan Rubenstein, ´ Ruoxin Sang, Dan Hurt, Gamaleldin Elsayed, Renshen Wang, Dave Lacey, Anastasija Ilic, Yao Zhao, ´ Adam Iwanicki, Alejandro Lince, Alexander Chen, Christina Lyu, Carl Lebsack, Jordan Griffith, Meenu Gaba, Paramjit Sandhu, Phil Chen, Anna Koop, Ravi Rajwar, Soheil Hassas Yeganeh, Solomon Chang, Rui Zhu, Soroush Radpour, Elnaz Davoodi, Ving Ian Lei, Yang Xu, Daniel Toyama, Constant Segal, Martin Wicke, Hanzhao Lin, Anna Bulanova, Adrià Puigdomènech Badia, Nemanja Rakicevi ´ c, Pablo Sprech- ´ mann, Angelos Filos, Shaobo Hou, Víctor Campos, Nora Kassner, Devendra Sachan, Meire Fortunato, Chimezie Iwuanyanwu, Vitaly Nikolaev, Balaji Lakshminarayanan, Sadegh Jazayeri, Mani Varadarajan, Chetan Tekur, Doug Fritz, Misha Khalman, David Reitter, Kingshuk Dasgupta, Shourya Sarcar, Tina Ornduff, Javier Snaider, Fantine Huot, Johnson Jia, Rupert Kemp, Nejc Trdin, Anitha Vijayakumar, Lucy Kim, Christof Angermueller, Li Lao, Tianqi Liu, Haibin Zhang, David Engel, Somer Greene, Anaïs White, Jessica Austin, Lilly Taylor, Shereen Ashraf, Dangyi Liu, Maria Georgaki, Irene Cai, Yana Kulizhskaya, Sonam Goenka, Brennan Saeta, Ying Xu, Christian Frank, Dario de Cesare, Brona Robenek, Harry Richardson, Mahmoud Alnahlawi, Christopher Yew, Priya Ponnapalli, Marco Tagliasacchi, Alex Korchemniy, Yelin Kim, Dinghua Li, Bill Rosgen, Kyle Levin, Jeremy Wiesner, Praseem Banzal, Praveen Srinivasan, Hongkun Yu, Çaglar Ünlü, David ˘ Reid, Zora Tung, Daniel Finchelstein, Ravin Kumar, Andre Elisseeff, Jin Huang, Ming Zhang, Ricardo Aguilar, Mai Giménez, Jiawei Xia, Olivier Dousse, Willi Gierke, Damion Yates, Komal Jalan, Lu Li, Eri Latorre-Chimoto, Duc Dung Nguyen, Ken Durden, Praveen Kallakuri, Yaxin Liu, Matthew Johnson, Tomy Tsai, Alice Talbert, Jasmine Liu, Alexander Neitz, Chen Elkind, Marco Selvi, Mimi Jasarevic, Livio Baldini Soares, Albert Cui, Pidong Wang, Alek Wenjiao Wang, Xinyu Ye, Krystal Kallarackal, Lucia Loher, Hoi Lam, Josef Broder, Dan HoltmannRice, Nina Martin, Bramandia Ramadhana, Mrinal Shukla, Sujoy Basu, Abhi Mohan, Nick Fernando, Noah Fiedel, Kim Paterson, Hui Li, Ankush Garg, Jane Park, DongHyun Choi, Diane Wu, Sankalp Singh, Zhishuai Zhang, Amir Globerson, Lily Yu, John Carpenter, Félix de Chaumont Quitry, Carey Radebaugh, Chu-Cheng Lin, Alex Tudor, Prakash Shroff, Drew Garmon, Dayou Du, Neera Vats, Han Lu, Shariq Iqbal, Alex Yakubovich, Nilesh Tripuraneni, James Manyika, Haroon Qureshi, Nan Hua, Christel Ngani, Maria Abi Raad, Hannah Forbes, Jeff Stanway, Mukund Sundararajan, Victor Ungureanu, Colton Bishop, Yunjie Li, Balaji Venkatraman, Bo Li, Chloe Thornton, Salvatore Scellato, Nishesh Gupta, Yicheng Wang, Ian Tenney, Xihui Wu, Ashish Shenoy, Gabriel Carvajal, Diana Gage Wright, Ben Bariach, Zhuyun Xiao, Peter Hawkins, Sid Dalmia, Clement Farabet, Pedro Valenzuela, Quan Yuan, Ananth Agarwal, Mia Chen, Wooyeol Kim, Brice Hulse, Nandita Dukkipati, Adam Paszke, Andrew Bolt, Kiam Choo, Jennifer Beattie, Jennifer Prendki, Harsha Vashisht, Rebeca SantamariaFernandez, Luis C. Cobo, Jarek Wilkiewicz, David Madras, Ali Elqursh, Grant Uy, Kevin Ramirez, Matt Harvey, Tyler Liechty, Heiga Zen, Jeff Seibert, Clara Huiyi Hu, Andrey Khorlin, Maigo Le, Asaf Aharoni, Megan Li, Lily Wang, Sandeep Kumar, Norman Casagrande, Jay Hoover, Dalia El Badawy, David Soergel, Denis Vnukov, Matt Miecnikowski, Jiri Simsa, Praveen Kumar, Thibault Sellam, Daniel Vlasic, Samira Daruki, Nir Shabat, John Zhang, Guolong Su, Jiageng Zhang, Jeremiah Liu, Yi Sun, Evan Palmer, Alireza Ghaffarkhah, Xi Xiong, Victor Cotruta, Michael Fink, Lucas Dixon, Ashwin Sreevatsa, Adrian Goedeckemeyer, Alek Dimitriev, Mohsen Jafari, Remi Crocker, Nicholas FitzGerald, Aviral Kumar, Sanjay Ghemawat, Ivan Philips, Frederick Liu, Yannie Liang, Rachel Sterneck, Alena Repina, Marcus Wu, Laura Knight, Marin Georgiev, Hyo Lee, Harry Askham, Abhishek Chakladar, Annie Louis, Carl Crous, Hardie Cate, Dessie Petrova, Michael Quinn, Denese Owusu-Afriyie, Achintya Singhal, Nan Wei, Solomon Kim, Damien Vincent, Milad Nasr, Christopher A. Choquette-Choo, Reiko Tojo, Shawn Lu, Diego de Las Casas, Yuchung Cheng, Tolga Bolukbasi, Katherine Lee, Saaber Fatehi, Rajagopal Ananthanarayanan, Miteyan Patel, Charbel Kaed, Jing Li, Shreyas Rammohan Belle, Zhe Chen, Jaclyn Konzelmann, Siim Põder, Roopal Garg, Vinod Koverkathu, Adam Brown, Chris Dyer, Rosanne Liu, Azade Nova, Jun Xu, Alanna Walton, Alicia Parrish, Mark Epstein, Sara McCarthy, Slav Petrov, Demis Hassabis, Koray Kavukcuoglu, Jeffrey Dean, and Oriol Vinyals. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.
  70. 70.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971.
  71. 71.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. 2023. Cogvlm: Visual expert for pretrained language models. ArXiv, abs/2311.03079.
  72. 72.Chen Henry Wu and Fernando De la Torre. 2023. A latent space of stochastic diffusion models for zeroshot image editing and guidance. In ICCV.
  73. 73.Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, and Lei Li. 2023. Instructscore: Towards explainable text generation evaluation with automatic feedback. ArXiv, abs/2305.14282.
  74. 74.Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023. The dawn of lmms: Preliminary explorations with gpt-4v(ision).
  75. 75.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. ArXiv, abs/2306.13549.
  76. 76.Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. 2023a. Magicbrush: A manually annotated dataset for instruction-guided image editing. NeurIPS dataset and benchmark track.
  77. 77.Lvmin Zhang and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:2302.05543.
  78. 78.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR.
  79. 79.Xinlu Zhang, Yujie Lu, Weizhi Wang, An Yan, Jun Yan, Lianke Qin, Heng Wang, Xifeng Yan, William Yang Wang, and Linda Ruth Petzold. 2023b. Gpt-4v(ision) as a generalist evaluator for vision-language tasks. ArXiv, abs/2311.01361.
  80. 80.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.

Citation

MLA
Ku, M., et al. “VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 12268–90, https://doi.org/10.18653/v1/2024.acl-long.663.
APA
Ku, M., Jiang, D., Wei, C., Yue, X., & Chen, W. (2024). VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12268–12290. https://doi.org/10.18653/v1/2024.acl-long.663
Chicago
Ku, M., D. Jiang, C. Wei, X. Yue, and W. Chen. 2024. “VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12268–90. https://doi.org/10.18653/v1/2024.acl-long.663.
Harvard
Ku, M. et al. (2024) “VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12268–12290. Available at: https://doi.org/10.18653/v1/2024.acl-long.663.
Vancouver
1. Ku M, Jiang D, Wei C, Yue X, Chen W (2024) VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12268–12290

BibTeX

@inproceedings{ku-etal-2024-viescore,
    title = "{VIES}core: Towards Explainable Metrics for Conditional Image Synthesis Evaluation",
    author = "Ku, Max  and
      Jiang, Dongfu  and
      Wei, Cong  and
      Yue, Xiang  and
      Chen, Wenhu",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.663/",
    doi = "10.18653/v1/2024.acl-long.663",
    pages = "12268--12290"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/