Text-Image Alignment for Diffusion-Based Perception

Neehar KondapaneniMarkus MarksManuel KnottRogério GuimarãesPietro Perona

article2024CVPR64 citations

Proposes an automated image-captioning prompting framework that improves text-image alignment in pretrained diffusion backbones to achieve state-of-the-art performance across semantic segmentation, depth estimation, and cross-domain visual tasks.

Listen

Diffusion models, widely celebrated for generating high-fidelity images from text, possess rich internal representations that make them promising backbones for discriminative computer vision tasks such as semantic segmentation, depth estimation, and object detection. However, standard methods for feeding text prompts into these generative architectures have relied on generic, unaligned inputs, such as averaging class embeddings across entire datasets. This introduces semantic noise and degrades feature representations, leaving the optimal strategy for prompting diffusion models in perception workflows an open and critical question.

The article demonstrates that aligning text prompts directly with input images and target domains significantly enhances the perceptual capabilities of diffusion backbones. It introduces Text-Aligned Diffusion Perception (TADP), an automated framework evaluating how automated captioning, latent scaling, and domain personalization improve visual task accuracy in both single-domain and cross-domain environments.

To establish these capabilities, the researchers conducted extensive empirical evaluations using Stable Diffusion backbones modified with task-specific decoding heads. The method automatically generates image captions using an off-the-shelf captioning model (BLIP-2) to ensure precise text-image alignment. In cross-domain settings—such as adapting models from daytime to nighttime driving or from photographic objects to artistic renderings—the researchers incorporated domain information into prompts using text modifiers and model personalization techniques (Textual Inversion and DreamBooth). The evaluation benchmarked performance across standard datasets including ADE20K, Pascal VOC, NYUv2, Cityscapes, Dark Zurich, and Watercolor2K.

The findings show that text-image alignment substantially boosts perception accuracy. On single-domain tasks, combining automated captioning with latent feature normalization improved semantic segmentation by approximately 4.0 mean Intersection over Union (mIoU) on Pascal VOC and 1.7 mIoU on ADE20K, while reducing depth estimation error on NYUv2 by 8% relative root mean square error (RMSE), setting a new state-of-the-art. Oracle experiments revealed that diffusion models are particularly sensitive to missing object classes (low recall) rather than extraneous classes. In cross-domain transfers, diffusion models demonstrated strong baseline generalization, and appending domain-specific text alignment or personalized tokens delivered top-tier results, reaching 72.2 AP50 on Watercolor2K and 60.8 mIoU on Nighttime Driving.

These results indicate that diffusion models do not automatically extract optimal semantic feature maps without targeted text guidance. By leveraging automated captioners, organizations can unlock superior visual perception performance without manually creating task-specific text descriptions or annotating massive target-domain datasets. This reduces annotation costs and accelerates deployment in domain-shifted environments like adverse weather navigation or non-photorealistic visual analysis.

Organizations developing perception systems should adopt automated captioning pipelines and latent scaling as drop-in upgrades for diffusion backbones. For domain adaptation projects, teams should use lightweight personalization methods such as Textual Inversion to align textual concepts with target visual styles without requiring heavy target-domain annotations. Further research should focus on refining specialized, high-recall captioners to approach theoretical upper-bound performance across diverse operating conditions.

The conclusions are supported by comprehensive benchmark comparisons, though confidence in cross-domain text alignment depends partially on the availability of target-domain style data or descriptive text prompts. Readers should note that while performance gains are substantial, the computational overhead of running automated captioners alongside diffusion backbones represents an operational trade-off in latency-critical production systems.

arXiv: 2310.00031
Cover for Text-Image Alignment for Diffusion-Based Perception

Abstract

Diffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusion-based semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our cross-domain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Diffusion models for single-domain vision tasks
  • 2.2. Image captioning
  • 2.3. Diffusion models for cross-domain vision tasks
  • 2.4. Cross-domain object detection
  • 3. Methods
  • 3.1. Text-Aligned Diffusion Perception (TADP)
  • 4. Results
  • 4.1. Latent scaling
  • 4.2. Single-domain alignment
  • 4.3. Cross-domain alignment
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Text-Aligned Diffusion Perception

    model/method

    Text-Aligned Diffusion Perception (TADP) adapts diffusion-pretrained vision models by replacing generic or class-averaged prompts with captions aligned to each input image. For an image xx, a captioner GG produces a caption y~\tilde y, which is encoded by the diffusion model’s text encoder into a conditioning representation CC:

    GTADP(x)=BLIP-2(x)=y~↦CTADP(x)G_{\mathrm{TADP}}(x)=\mathrm{BLIP\text{-}2}(x)=\tilde y\mapsto C_{\mathrm{TADP}}(x).

    In the single-domain setting, C=CTADP(x)C=C_{\mathrm{TADP}}(x). In the cross-domain setting, PP denotes broad information about the target domain, such as watercolor, comic, foggy, or nighttime style; a caption modifier MM produces a target-domain text suffix or prefix M(P)sM(P)_s, and the conditioning becomes C=CTADP(x)⊕M(P)sC=C_{\mathrm{TADP}}(x)\oplus M(P)_s, where ⊕\oplus denotes text concatenation. The modifier may also alter the diffusion backbone, yielding θϵ′=M(P)ϵ(θϵ)\theta'_\epsilon=M(P)_\epsilon(\theta_\epsilon). TADP therefore aligns the diffusion representation both with the image content and, when needed, with the target-domain appearance.

  2. Knowl 2 — Diffusion feature extraction with latent scaling

    model/method

    TADP uses a one-step diffusion feature extractor derived from a diffusion-pretrained perception backbone. For an input image xx, the image encoder EE produces a latent z0=E(x)z_0=E(x), which is scaled by the fixed Stable Diffusion factor 0.182150.18215: zˉ0=0.18215 z0\bar z_0=0.18215\,z_0. The scaled latent and text conditioning CC are passed through the denoising U-Net ϵθ\epsilon_\theta at timestep t=0t=0. The U-Net supplies cross-attention maps AA and multi-scale feature maps FF; these are concatenated into V=A⊕FV=A\oplus F and decoded by a task-specific head HH:

    p^=H(V)=H ⁣(A⊕F).\hat p=H(V)=H\!\left(A\oplus F\right).

    Here p^\hat p is the prediction for semantic segmentation, object detection, or monocular depth estimation. The diffusion backbone and task head are trained with the task-specific loss while the text conditioning controls the cross-attention maps and the resulting feature representation. Latent scaling alone improves the class-averaged-prompt baseline by approximately 0.80.8 Pascal VOC mIoU, 0.30.3 ADE20K mIoU, and 5.5%5.5\% relative NYUv2 RMSE.

  3. Knowl 3 — Why class-averaged EOS prompts are inadequate

    model/method

    Earlier diffusion-perception systems represent each dataset class by averaging the first end-of-sentence (EOS) token from 80 CLIP sentence templates. For a class set BB and CLIP token dimension 768768, this produces Cavg∈R∣B∣×768C_{\mathrm{avg}}\in\mathbb{R}^{|B|\times768}. TADP finds that this averaging is poorly suited to diffusion cross-attention: although it can be useful for measuring similarity in CLIP space, it degrades the spatial attention maps used by the diffusion backbone.

    Two alternatives are evaluated. The first passes the CLIP token sequence for each class name directly to cross-attention:

    GClassEmbs(B)=concat⁡ ⁣(CLIP(b)∣b∈B)↦CClassEmbs.G_{\mathrm{ClassEmbs}}(B)=\operatorname{concat}\!\left(\mathrm{CLIP}(b)\mid b\in B\right)\mapsto C_{\mathrm{ClassEmbs}}.

    The second forms a grammatical-free string containing the class names separated by spaces and tokenizes that string:

    GClassNames(B)=concat⁡ ⁣(⟨space⟩+b∣b∈B)↦CClassNames.G_{\mathrm{ClassNames}}(B)=\operatorname{concat}\!\left(\langle\mathrm{space}\rangle+b\mid b\in B\right)\mapsto C_{\mathrm{ClassNames}}.

    On Pascal VOC2012 segmentation, direct class-token embeddings achieve 82.7282.72 mIoU, below the latent-scaled averaged-EOS baseline of 83.0683.06, whereas the space-separated class-name prompt achieves 84.0884.08. The attention analysis also shows that absent classes can receive highly localized attention, so merely listing every dataset class does not reliably align the representation with the current image.

  4. Knowl 4 — Automated BLIP-2 captions improve single-domain alignment

    empirical result

    TADP uses BLIP-2 to generate an image caption for every image in Pascal VOC2012, ADE20K, and NYUv2, then feeds the caption tokens directly into the diffusion backbone. Compared with the latent-scaled VPD baseline, image captions improve Pascal segmentation by approximately 44 mIoU, ADE20K segmentation by approximately 1.41.4 mIoU, and NYUv2 depth estimation by approximately 4%4\% relative RMSE. The gains are larger when the task-specific model is trained briefly: ADE20K improves by approximately 55 mIoU after 4k iterations and 2.42.4 mIoU after 8k iterations; the one-epoch NYUv2 model improves by approximately 2.4%2.4\% relative RMSE.

    The captioner is task-agnostic and open-vocabulary, so it can identify image content without requiring a fixed list of segmentation or detection classes. This alignment is more effective than providing all possible class names, even though the latter contains the complete dataset vocabulary.

  5. Knowl 5 — Single-domain benchmark performance

    data/table

    The reported single-domain experiments compare diffusion-perception prompting strategies on Pascal VOC2012 semantic segmentation, ADE20K semantic segmentation, and NYUv2 monocular depth estimation. Pascal and ADE20K use mIoU, with Pascal reporting single-scale mIoU and ADE20K reporting single-scale and multi-scale mIoU; NYUv2 uses RMSE, for which lower values are better. TADP-40 denotes BLIP-2 captions constrained to have at least 40 tokens.

    Could not parse LaTeX table

    TADP-40 is the best practical captioning configuration in these experiments. It establishes the best result among the compared diffusion-pretrained ADE20K models, reaches 87.1187.11 mIoU on Pascal, and reduces default-schedule NYUv2 RMSE from 0.2350.235 with latent scaling alone to 0.2250.225. The oracle results, which use ground-truth object names and are not available at deployment, indicate substantially greater potential from perfectly aligned prompts.

  6. Knowl 6 — Caption content, length, and oracle alignment

    empirical result

    TADP’s caption ablations show that caption content matters more than grammatical completeness. Increasing the BLIP-2 minimum caption length from 0 to 20 or 40 tokens gives small but consistent gains; 40-token captions provide an average relative improvement of approximately 0.75%0.75\% over zero-token-minimum captions, with gains of about 0.80.8 Pascal mIoU, 0.90.9 ADE20K mIoU, and 0.6%0.6\% relative NYUv2 RMSE. Filtering captions to nouns only produces 86.3586.35 mIoU on Pascal, essentially matching the 86.1986.19 mIoU of the full 20-token caption, suggesting that object nouns carry most of the useful alignment signal.

    For an oracle analysis, B(x)B(x) denotes the set of ground-truth object classes present in image xx. The oracle caption consists only of those class names:

    GOracle(x)=concat⁡ ⁣(⟨space⟩+b∣b∈B(x))↦COracle(x).G_{\mathrm{Oracle}}(x)=\operatorname{concat}\!\left(\langle\mathrm{space}\rangle+b\mid b\in B(x)\right)\mapsto C_{\mathrm{Oracle}}(x).

    This oracle reaches 89.8589.85 Pascal VOC2012 mIoU and approximately 7272 ADE20K mIoU, substantially above automatic captioning. ADE20K perturbations show that missing classes are more damaging than extra classes: with caption recall fixed at 0.500.50, increasing precision has less than a 11 mIoU-point effect, while recall reductions produce approximately 66, 99, and 1313 mIoU-point losses at precision levels 0.500.50, 0.750.75, and 1.001.00, respectively. At higher recall, precision also matters, causing approximately 33 and 77 mIoU-point losses at recall 0.750.75 and 1.001.00. BLIP-2 captions behave approximately like prompts with perfect precision and recall around 0.500.50.

  7. Knowl 7 — Cross-domain training with target-domain caption modifiers

    model/method

    For cross-domain transfer, TADP trains on labeled source-domain images and evaluates on a different target domain. BLIP-2 first supplies image-aligned source captions. A target-domain modifier is then applied consistently during training and testing. The null modifier adds no target-style information. The simple modifier appends a hand-written style phrase, such as a watercolor, comic, dark-night, or foggy-photo description. Textual Inversion learns a new target-style token from unlabeled target-domain images and inserts it into the caption. DreamBooth learns a target-style token from the same kind of unlabeled target images and additionally fine-tunes the Stable Diffusion backbone.

    The cross-domain experiments use Pascal VOC as the source for Watercolor2K and Comic2K object detection, and Cityscapes as the source for Dark Zurich-val and Nighttime Driving semantic segmentation. Textual Inversion and DreamBooth use only unlabeled target-domain image examples for personalization; they do not require target labels. At test time, GPT-3.5 is used to remove target-style words that BLIP-2 may have spontaneously inserted into target-image captions, so the captions receive exactly the intended modifier rather than redundant or inconsistent domain information.

  8. Knowl 8 — Cross-domain transfer results

    data/table

    The cross-domain results compare target-domain caption modifiers using AP and AP50 for object detection and mIoU for semantic segmentation. All TADP variants use image-aligned captions; they differ only in how target-domain information is supplied. The null variant demonstrates the strength of diffusion pretraining without explicit target-style text, while the personalized variants test whether unlabeled target images can improve alignment.

    Could not parse LaTeX table

    The method obtains the reported best result on Watercolor2K with 72.272.2 AP50 and the best reported Nighttime Driving result with 60.860.8 mIoU from Textual Inversion. The null configuration already performs strongly without target-style text, achieving 42.842.8 Dark Zurich-val mIoU and 72.172.1 Watercolor2K AP50. On Comic2K, Textual Inversion gives the highest TADP AP at 33.233.2, although the strongest external method in the comparison uses extra target-domain training data.

  9. Knowl 9 — Personalization is more reliable than hand-written target styles

    empirical result

    Averaged over the two cross-domain segmentation datasets, the null, simple, Textual Inversion, and DreamBooth modifiers obtain respectively 50.550.5, 48.048.0, 51.151.1, and 49.749.7 average mIoU. Averaged over the two cross-domain detection datasets, the same modifiers obtain 36.636.6, 37.737.7, 38.238.2, and 38.138.1 average AP. Thus, Textual Inversion is the strongest overall modifier in the reported averages, while DreamBooth is close behind.

    The result is not explained merely by adding any style phrase. Control modifiers describing a nearby but incorrect domain and an unrelated domain reduce performance across the transfer datasets. For example, nearby styles reach 41.941.9 and 56.956.9 mIoU on Dark Zurich-val and Nighttime Driving, while unrelated styles reach 42.342.3 and 55.155.1; the corresponding object-detection AP values are 31.831.8 and 32.032.0 on Comic2K. The experiments therefore support target-domain alignment only when the textual or personalized style matches the actual target distribution. Model personalization is particularly useful when a target appearance cannot be represented accurately by words alone.

  10. Knowl 10 — Limitations and unresolved alignment problems

    limitation

    TADP remains sensitive to caption alignment. The oracle experiments show a large gap between automatic captions and captions containing the complete set of objects in an image, especially because missed classes are harmful. A hand-written target-style phrase can also degrade cross-domain performance when it is inaccurate, and the paper notes that target-domain appearance is sometimes difficult to express with text alone. Textual Inversion and DreamBooth reduce this problem using unlabeled target images, but they require additional personalization data and procedures. The authors identify extension to multiple unseen target domains and more task-specific, closed-vocabulary captioners as unresolved directions.

Coverage note — No substantial contributed component was omitted; background, related work, acknowledgements, and reference material were excluded as non-contributory.

References

  1. 1.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Ji aming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers. arXiv preprint arXiv:2211.01324, 2022. 2
  2. 2.Yasser Benigmim, Subhankar Roy, Slim Essid, Vicky Kalogeiton, and Stephane Lathuili ere. One-shot Unsupervised Domain Adaptation with Personalized Diffusion Models. arXiv preprint arXiv:2303.18080, 2023. 2
  3. 3.Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Muller. ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth. arXiv preprint arXiv:2302.12288, 2023. 5
  4. 4.Steven Bird, Ewan Klein, and Edward Loper. Natural language processing with Python: analyzing text with the natural language toolkit. O’Reilly Media, Inc., 2009. 6
  5. 5.Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, and Kristian Kersting. SEGA: Instructing Diffusion using Semantic Dimensions. arXiv preprint arXiv:2301.12247, 2023. 1, 2
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 7
  7. 7.David Bruggemann, Christos Sakaridis, Prune Truong, and Luc Van Gool. Refign: Align and Refine for Adaptation of Semantic Segmentation to Adverse Conditions. 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022. 7, 8
  8. 8.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Y. Qiao. Vision Transformer Adapter for Dense Predictions. arXiv preprint arXiv:2205.08534, 2022. 5
  9. 9.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3213–3223, 2016. 2, 7
  10. 10.Dengxin Dai and Luc Van Gool. Dark Model Adaptation: Semantic Image Segmentation from Daytime to Nighttime. 2018 21st International Conference on Intelligent Transportation Systems (ITSC), pages 3819–3824, 2018. 2, 7
  11. 11.Jinhong Deng, Wen Li, Yuhua Chen, and Lixin Duan. Unbiased mean teacher for cross-domain object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4091–4101, 2021. 3, 7
  12. 12.David Eigen, Christian Puhrsch, and Rob Fergus. Depth Map Prediction from a Single Image using a Multi-Scale Deep Network. In Advances in Neural Information Processing Systems 27 (NIPS 2014), 2014. 14
  13. 13.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html. 2, 7
  14. 14.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2012 (VOC2012), 2012. 2, 7
  15. 15.Yuxin Fang, Wen Wang, Binhui Xie, Quan-Sen Sun, Ledell Yu Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19358–19369, 2022. 5
  16. 16.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. arXiv preprint arXiv:2208.01618, 2022. 2, 7, 8
  17. 17.Rui Gong, Martin Danelljan, Han Sun, Julio Delgado Mangas, and Luc Van Gool. Prompting Diffusion Representations for Cross-Domain Semantic Segmentation. arXiv preprint arXiv:2307.02138, 2023. 1, 2, 3, 7
  18. 18.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv preprint arXiv:2208.01626, 2022. 2
  19. 19.Luwei Hou, Yu Zhang, Kui Fu, and Jia Li. Informative and consistent correspondence mining for cross-domain weakly supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9929–9938, 2021. 3, 7
  20. 20.Lukas Hoyer, Dengxin Dai, and Luc Van Gool. DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9914–9925, 2022. 7
  21. 21.Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, and Kiyoharu Aizawa. Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5001–5009, 2018. 2, 3, 7
  22. 22.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. arXiv preprint arXiv:2102.05918, 2021. 1
  23. 23.Junguang Jiang, Baixu Chen, Jianmin Wang, and Mingsheng Long. Decoupled adaptation for cross-domain object detection. arXiv preprint arXiv:2110.02578, 2021. 3
  24. 24.Alexander Kirillov, Ross B. Girshick, Kaiming He, and Piotr Dollar. Panoptic Feature Pyramid Networks. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6392–6401, 2019. 8, 14
  25. 25.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. arXiv preprint arXiv:2301.12597, 2023. 2, 4, 5
  26. 26.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm. arXiv preprint arXiv:2110.05208, 2022. 1
  27. 27.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. Swin Transformer V2: Scaling Up Capacity and Resolution. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11999–12009, 2021. 5
  28. 28.Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell. Diffusion hyperfeatures: Searching through time and space for semantic correspondence. arXiv preprint arXiv:2305.14334, 2023. 1, 2, 3
  29. 29.Jia Ning, Chen Li, Zheng Zhang, Zigang Geng, Qi Dai, Kun He, and Han Hu. All in Tokens: Unifying Output Space of Visual Tasks via Soft Token. arXiv preprint arXiv:2301.02229, 2023. 5
  30. 30.Shengxiong Ouyang, Xinglu Wang, Kejie Lyu, and Yingming Li. Pseudo-label generation-evaluation framework for cross domain weakly supervised object detection. In 2021 IEEE International Conference on Image Processing (ICIP), pages 724–728. IEEE, 2021. 3, 7
  31. 31.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. International Conference on Machine Learning, pages 8748–8763, 2021. 1, 2, 3
  32. 32.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv preprint arXiv:2204.06125, 2022. 1, 2
  33. 33.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18061–18070, 2022. 5
  34. 34.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1137–1149, 2017. 14
  35. 35.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, 2022. 1, 2, 3, 4
  36. 36.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, pages 234–241. Springer International Publishing, Cham, 2015. 1
  37. 37.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. arXiv preprint arXiv:2208.12242, 2022. 2, 7, 8
  38. 38.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv preprint arXiv:2205.11487, 2022. 1, 2
  39. 39.Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Guided Curriculum Model Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7373–7382, 2019. 2, 7
  40. 40.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. arXiv preprint arXiv:2111.02114, 2021. 3
  41. 41.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022. 2
  42. 42.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor Segmentation and Support Inference from RGBD Images. European Conference on Computer Vision (ECCV), 2012. 2
  43. 43.Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. arXiv preprint arXiv:2306.03881, 2023. 1, 2, 3
  44. 44.Barıs¸ Batuhan Topal, Deniz Yuret, and Tevfik Metin Sezgin. Domain-adaptive self-supervised pre-training for face & body detection in drawings. arXiv preprint arXiv:2211.10641, 2022. 3, 7, 8
  45. 45.Eric Tzeng, Judy Hoffman, Kate Saenko, and Trevor Darrell. Adversarial discriminative domain adaptation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7167–7176, 2017. 7
  46. 46.Vidit Vidit, Martin Engilberge, and Mathieu Salzmann. CLIP the Gap: A Single Domain Generalization Approach for Object Detection. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3219–3229, 2023. 3, 7
  47. 47.Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, and Chang Zhou. ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities. arXiv preprint arXiv:2305.11172, 2023. 5
  48. 48.Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiao-hua Hu, Tong Lu, Lewei Lu, Hongsheng Li, Xiaogang Wang, and Y. Qiao. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14408–14419, 2022. 5
  49. 49.Wen Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiangbo Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19175–19186, 2023. 5
  50. 50.Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, and Chunhua Shen. DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models. arXiv preprint arXiv:2303.11681, 2023. 2
  51. 51.Yunqiu Xu, Yifan Sun, Zongxin Yang, Jiaxu Miao, and Yi Yang. H2fa r-cnn: Holistic and hierarchical feature alignment for cross-domain weakly supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14329–14339, 2022. 3, 7
  52. 52.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Benton C. Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. arXiv preprint arXiv:2206.10789, 2022. 1, 2
  53. 53.Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou, and Jiwen Lu. Unleashing Text-to-Image Diffusion Models for Visual Perception. arXiv preprint arXiv:2303.02153, 2023. 1, 2, 3, 4, 5, 14
  54. 54.Zhen Zhao, Yuhong Guo, Haifeng Shen, and Jieping Ye. Adaptive object detection with dual multi-label prediction. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16, pages 54–69. Springer, 2020. 3, 7
  55. 55.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene Parsing through ADE20K Dataset. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5122–5130, 2017. 2

Citation

MLA
Kondapaneni, N., et al. “Text-image Alignment for Diffusion-based Perception”. arXiv, 2023, http://arxiv.org/abs/2310.00031v3.
APA
Kondapaneni, N., Marks, M., Knott, M., Guimaraes, R., & Perona, P. (2023). Text-image Alignment for Diffusion-based Perception. arXiv. http://arxiv.org/abs/2310.00031v3
Chicago
Kondapaneni, N., M. Marks, M. Knott, R. Guimaraes, and P. Perona. 2023. “Text-image Alignment for Diffusion-based Perception”. arXiv. http://arxiv.org/abs/2310.00031v3.
Harvard
Kondapaneni, N. et al. (2023) “Text-image Alignment for Diffusion-based Perception”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.00031v3.
Vancouver
1. Kondapaneni N, Marks M, Knott M, Guimaraes R, Perona P (2023) Text-image Alignment for Diffusion-based Perception. arXiv

BibTeX

@article{kondapaneni2023text,
  title = {Text-image Alignment for Diffusion-based Perception},
  author = {Kondapaneni, Neehar and Marks, Markus and Knott, Manuel and Guimaraes, Rogerio and Perona, Pietro},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.00031v3},
  eprint = {2310.00031}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE