DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback
Jiao SunDeqing FuYushi HuSu WangRoyi RassinDa-Cheng JuanDana AlonCharles HerrmannSjoerd van SteenkisteRanjay Krishna
Presents DreamSync, an iterative self-training framework that uses vision-language models for automated text-alignment and aesthetic scoring to fine-tune text-to-image diffusion models with LoRA without requiring human labels or reinforcement learning.
Modern text-to-image models often struggle to generate images that both faithfully reflect user prompts and maintain high visual appeal. Common failure points include mishandling complex multi-object compositions, misattributing colors or traits, and incorrectly rendering visual text. While existing solutions attempt to address these shortcomings through inference adjustments or reinforcement learning using human feedback, they typically require expensive human annotation, alter underlying architectures, or degrade the overall aesthetic quality of the images.
The article demonstrates an automated self-improving training framework called DreamSync that enhances text faithfulness without sacrificing visual appeal. The primary objective is to evaluate how vision-language models—specifically those for visual question answering and aesthetic quality scoring—can provide autonomous feedback to iteratively refine text-to-image diffusion models without needing labeled human data.
The DreamSync approach functions via an automated four-step loop consisting of sampling, evaluating, filtering, and fine-tuning. First, a large language model automatically generates diverse synthetic prompts and associated question-answer pairs. The generative image model produces several candidate images per prompt. Vision-language models then score each image for text alignment and visual quality. Candidates exceeding strict quality thresholds are selected to fine-tune the generative model using parameter-efficient Low-Rank Adaptation. The updated model subsequently repeats the cycle across multiple iterations.
The evaluation yields several key findings across automated benchmarks and human evaluations. When applied to Stable Diffusion XL, DreamSync improves text alignment scores by 1.7 percentage points on the TIFA benchmark and 2.9 percentage points on the DSG1K benchmark, while simultaneously raising aesthetic scores by 3.4 points. In human evaluations covering over 1,000 prompts, DreamSync consistently outperforms baseline models across all semantic categories, achieving its largest individual gain of 18.52% in text rendering accuracy. Furthermore, on older base architectures like Stable Diffusion v1.4, DreamSync outperforms both training-free and reinforcement-learning-based alignment baselines in text faithfulness and human preference reward scores without introducing stylistic distortions.
These findings indicate that automated multimodal feedback is a cost-effective, scalable substitute for expensive human-in-the-loop annotations. Organizations deploying generative image models can achieve significant performance gains in instruction-following reliability while lowering data curation costs and mitigating aesthetic degradation. Moreover, the methodology generalizes well across both in-distribution and out-of-distribution prompts.
Decision-makers should consider adopting automated multimodal self-training pipelines as a low-cost, parameter-efficient method to enhance image generation capabilities. For future implementations, organizations can explore grounding feedback mechanisms to provide localized, spatial annotations or integrating dynamic prompt generation for continuous model improvement. However, stakeholders should note that the method's ultimate performance remains constrained by the underlying base model, as complex multi-object attribute binding and fine texture rendering still exhibit occasional failure modes.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). DreamSync fine-tunes Stable Diffusion XL, so understanding SDXL’s architecture and capabilities clarifies the base model the feedback loop adapts.
No sufficiently relevant recommendations were found.
