DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback

Jiao SunDeqing FuYushi HuSu WangRoyi RassinDa-Cheng JuanDana AlonCharles HerrmannSjoerd van SteenkisteRanjay Krishna

article2025NAACL70 citations

Presents DreamSync, an iterative self-training framework that uses vision-language models for automated text-alignment and aesthetic scoring to fine-tune text-to-image diffusion models with LoRA without requiring human labels or reinforcement learning.

Listen

Modern text-to-image models often struggle to generate images that both faithfully reflect user prompts and maintain high visual appeal. Common failure points include mishandling complex multi-object compositions, misattributing colors or traits, and incorrectly rendering visual text. While existing solutions attempt to address these shortcomings through inference adjustments or reinforcement learning using human feedback, they typically require expensive human annotation, alter underlying architectures, or degrade the overall aesthetic quality of the images.

The article demonstrates an automated self-improving training framework called DreamSync that enhances text faithfulness without sacrificing visual appeal. The primary objective is to evaluate how vision-language models—specifically those for visual question answering and aesthetic quality scoring—can provide autonomous feedback to iteratively refine text-to-image diffusion models without needing labeled human data.

The DreamSync approach functions via an automated four-step loop consisting of sampling, evaluating, filtering, and fine-tuning. First, a large language model automatically generates diverse synthetic prompts and associated question-answer pairs. The generative image model produces several candidate images per prompt. Vision-language models then score each image for text alignment and visual quality. Candidates exceeding strict quality thresholds are selected to fine-tune the generative model using parameter-efficient Low-Rank Adaptation. The updated model subsequently repeats the cycle across multiple iterations.

The evaluation yields several key findings across automated benchmarks and human evaluations. When applied to Stable Diffusion XL, DreamSync improves text alignment scores by 1.7 percentage points on the TIFA benchmark and 2.9 percentage points on the DSG1K benchmark, while simultaneously raising aesthetic scores by 3.4 points. In human evaluations covering over 1,000 prompts, DreamSync consistently outperforms baseline models across all semantic categories, achieving its largest individual gain of 18.52% in text rendering accuracy. Furthermore, on older base architectures like Stable Diffusion v1.4, DreamSync outperforms both training-free and reinforcement-learning-based alignment baselines in text faithfulness and human preference reward scores without introducing stylistic distortions.

These findings indicate that automated multimodal feedback is a cost-effective, scalable substitute for expensive human-in-the-loop annotations. Organizations deploying generative image models can achieve significant performance gains in instruction-following reliability while lowering data curation costs and mitigating aesthetic degradation. Moreover, the methodology generalizes well across both in-distribution and out-of-distribution prompts.

Decision-makers should consider adopting automated multimodal self-training pipelines as a low-cost, parameter-efficient method to enhance image generation capabilities. For future implementations, organizations can explore grounding feedback mechanisms to provide localized, spatial annotations or integrating dynamic prompt generation for continuous model improvement. However, stakeholders should note that the method's ultimate performance remains constrained by the underlying base model, as complex multi-object attribute binding and fine texture rendering still exhibit occasional failure modes.

Sun et al (2025).pdf

No sufficiently relevant recommendations were found.

Cover for DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback

Abstract

Despite their widespread success, Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text. We introduce DreamSync, a simple yet effective training algorithm that improves T2I models to be faithful to the text input. DreamSync utilizes large vision-language models (VLMs) to effectively identify the fine-grained discrepancies between generated images and the text inputs and enable T2I models to self-improve without labeled data. First, it prompts the model to generate several candidate images for a given input text. Then, it uses two VLMs to select the best generation: a Visual Question Answering model that measures the alignment of generated images to the text, and another that measures the generation’s aesthetic quality. After selection, we use LoRA to iteratively finetune the T2I model to guide its generation towards the selected best generations. DreamSync does not need any additional human annotation, model architecture changes, or reinforcement learning. Despite its simplicity, DreamSync improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic) and human evaluation shows that DreamSync improves text rendering compared to SDXL by 18.5% on DSG1K benchmark.

Citation

MLA
Sun, J., et al. “DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 5920–45, https://doi.org/10.18653/v1/2025.naacl-long.304.
APA
Sun, J., Fu, D., Hu, Y., Wang, S., Rassin, R., Juan, D.-C., Alon, D., Herrmann, C., Steenkiste, S. van ., Krishna, R., & Rashtchian, C. (2025). DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5920–5945. https://doi.org/10.18653/v1/2025.naacl-long.304
Chicago
Sun, J., D. Fu, Y. Hu, et al. 2025. “DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5920–45. https://doi.org/10.18653/v1/2025.naacl-long.304.
Harvard
Sun, J. et al. (2025) “DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5920–5945. Available at: https://doi.org/10.18653/v1/2025.naacl-long.304.
Vancouver
1. Sun J, Fu D, Hu Y, et al (2025) DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5920–5945

BibTeX

@inproceedings{sun-etal-2025-dreamsync,
    title = "{D}ream{S}ync: Aligning Text-to-Image Generation with Image Understanding Feedback",
    author = "Sun, Jiao  and
      Fu, Deqing  and
      Hu, Yushi  and
      Wang, Su  and
      Rassin, Royi  and
      Juan, Da-Cheng  and
      Alon, Dana  and
      Herrmann, Charles  and
      van Steenkiste, Sjoerd  and
      Krishna, Ranjay  and
      Rashtchian, Cyrus",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.304/",
    doi = "10.18653/v1/2025.naacl-long.304",
    pages = "5920--5945",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/