Text-Image Alignment for Diffusion-Based Perception
Neehar KondapaneniMarkus MarksManuel KnottRogério GuimarãesPietro Perona
Proposes an automated image-captioning prompting framework that improves text-image alignment in pretrained diffusion backbones to achieve state-of-the-art performance across semantic segmentation, depth estimation, and cross-domain visual tasks.
Diffusion models, widely celebrated for generating high-fidelity images from text, possess rich internal representations that make them promising backbones for discriminative computer vision tasks such as semantic segmentation, depth estimation, and object detection. However, standard methods for feeding text prompts into these generative architectures have relied on generic, unaligned inputs, such as averaging class embeddings across entire datasets. This introduces semantic noise and degrades feature representations, leaving the optimal strategy for prompting diffusion models in perception workflows an open and critical question.
The article demonstrates that aligning text prompts directly with input images and target domains significantly enhances the perceptual capabilities of diffusion backbones. It introduces Text-Aligned Diffusion Perception (TADP), an automated framework evaluating how automated captioning, latent scaling, and domain personalization improve visual task accuracy in both single-domain and cross-domain environments.
To establish these capabilities, the researchers conducted extensive empirical evaluations using Stable Diffusion backbones modified with task-specific decoding heads. The method automatically generates image captions using an off-the-shelf captioning model (BLIP-2) to ensure precise text-image alignment. In cross-domain settings—such as adapting models from daytime to nighttime driving or from photographic objects to artistic renderings—the researchers incorporated domain information into prompts using text modifiers and model personalization techniques (Textual Inversion and DreamBooth). The evaluation benchmarked performance across standard datasets including ADE20K, Pascal VOC, NYUv2, Cityscapes, Dark Zurich, and Watercolor2K.
The findings show that text-image alignment substantially boosts perception accuracy. On single-domain tasks, combining automated captioning with latent feature normalization improved semantic segmentation by approximately 4.0 mean Intersection over Union (mIoU) on Pascal VOC and 1.7 mIoU on ADE20K, while reducing depth estimation error on NYUv2 by 8% relative root mean square error (RMSE), setting a new state-of-the-art. Oracle experiments revealed that diffusion models are particularly sensitive to missing object classes (low recall) rather than extraneous classes. In cross-domain transfers, diffusion models demonstrated strong baseline generalization, and appending domain-specific text alignment or personalized tokens delivered top-tier results, reaching 72.2 AP50 on Watercolor2K and 60.8 mIoU on Nighttime Driving.
These results indicate that diffusion models do not automatically extract optimal semantic feature maps without targeted text guidance. By leveraging automated captioners, organizations can unlock superior visual perception performance without manually creating task-specific text descriptions or annotating massive target-domain datasets. This reduces annotation costs and accelerates deployment in domain-shifted environments like adverse weather navigation or non-photorealistic visual analysis.
Organizations developing perception systems should adopt automated captioning pipelines and latent scaling as drop-in upgrades for diffusion backbones. For domain adaptation projects, teams should use lightweight personalization methods such as Textual Inversion to align textual concepts with target visual styles without requiring heavy target-domain annotations. Further research should focus on refining specialized, high-recall captioners to approach theoretical upper-bound performance across diverse operating conditions.
The conclusions are supported by comprehensive benchmark comparisons, though confidence in cross-domain text alignment depends partially on the availability of target-domain style data or descriptive text prompts. Readers should note that while performance gains are substantial, the computational overhead of running automated captioners alongside diffusion backbones represents an operational trade-off in latency-critical production systems.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). It demonstrates how cross-attention maps in text-to-image diffusion models localize semantic concepts, providing the foundational interpretability framework utilized by the source paper to extract perceptual features.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). It introduces cross-attention control mechanisms that establish how text tokens spatially ground visual objects in diffusion models, which underpins prompt-aligned perception.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). It establishes parameter-efficient fine-tuning via cross-attention weight optimization, which directly informs the model personalization techniques applied in the source.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). It introduces spatial conditioning for diffusion backbones, motivating methods that leverage generative diffusion features for downstream spatial understanding tasks.
- Paper: Improving CLIP Training with Language Rewrites, Lijie Fan et al. (2023). It shows that augmenting and modifying textual descriptions significantly enhances visual-language alignment, setting the precedent for caption modification strategies in perception.
- Paper: Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing, Bingyan Liu et al. (2024). It provides a probing analysis of cross- and self-attention semantic structures in Stable Diffusion, extending the understanding of attention-level text-image alignment examined in the source.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). It builds on cross-attention spatial grounding concepts by developing specialized attention controllers to separate and localize multi-instance visual semantics.
- Paper: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, Sihyun Yu et al. (2025). It extends representation alignment in diffusion architectures by directly coupling internal transformer representations with self-supervised visual perceptual encoders.
- Paper: There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation, Gabe Guo et al. (2026). It formalizes bidirectional diffusion bridges for multimodal translation, expanding upon directional text-image alignment frameworks.
