GLIGEN: Open-Set Grounded Text-to-Image Generation

Yuheng LiHaotian LiuQingyang WuFangzhou MuJianwei YangJianfeng GaoChunyuan LiYong Jae Lee

article2023CVPR881 citations

Proposes GLIGEN, a framework that injects spatial grounding inputs into frozen pre-trained diffusion models via gated layers, enabling precise layout-controlled image generation across open-world concepts without retraining the base model.

Listen

Existing text-to-image diffusion models generate impressive visual content from open-ended prompts, but they frequently struggle with precise spatial controllability, object placement, and editing accuracy. In standard text-to-image workflows, users cannot reliably specify exact object locations, poses, or boundaries, which limits their utility in production environments that require fine-grained scene composition.

The article evaluates GLIGEN (Grounded-Language-to-Image Generation), an approach designed to endow text-to-image diffusion models with open-set grounded generation capabilities using bounding boxes, human keypoints, reference images, and spatial maps.

The authors integrated grounded condition tokens into pretrained diffusion backbones using newly added gated self-attention layers while keeping the core generative weights intact. The evaluation spanned multiple benchmark datasets (COCO, LVIS, GoldG, Objects365, CC3M, and SBU) and tested tasks including text-grounded inpainting, image-grounded editing, keypoint-guided human generation, and layout-to-image synthesis.

The experiments show that incorporating spatial grounding substantially improves alignment and image fidelity across tasks. In text-grounded inpainting, the method achieved higher precision across all object sizes compared to baseline latent diffusion, maintaining an average precision of 25.6% on large objects where the baseline dropped to 14.6%. For human keypoint conditioning on COCO, the approach reached an average precision of 31.8% and a Fréchet Inception Distance of 31.02, dramatically outperforming the dedicated translation baseline pix2pixHD (15.8% AP and 142.4 FID). Pretraining across large-scale datasets and finetuning on LVIS yielded an average precision of 14.9% and an FID of 6.25, significantly surpassing supervised layout models like LAMA. Furthermore, gated self-attention proved superior to cross-attention mechanisms by allowing necessary visual token interaction.

These results demonstrate that spatial grounding can be added efficiently to pretrained models without retraining foundational weights from scratch. This reduces computational training costs and operational risk while providing reliable layout control for commercial creative tools, precision editing, and synthetic data pipelines.

Organizations developing controllable generative media should adopt gated self-attention architectures and leverage mixed caption and detection datasets during pretraining. When deploying spatial conditioning, teams should pair coarse bounding boxes for category-agnostic layout control and reserve keypoint conditioning for specific humanoid poses.

A key limitation identified is that keypoint grounding does not generalize across non-humanoid object categories, as anatomical part associations are less transferable than generic bounding boxes. Confidence in the reported image quality and layout correspondence remains high across standard bounding box and human pose benchmarks.

Cover for GLIGEN: Open-Set Grounded Text-to-Image Generation

Abstract

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN, Grounded-Language-to-Image Generation, a novel approach that builds upon and extends the functionality of existing pre-trained text-to-image diffusion models by enabling them to also be conditioned on grounding inputs. To preserve the vast concept knowledge of the pre-trained model, we freeze all of its weights and inject the grounding information into new trainable layers via a gated mechanism. Our model achieves open-world grounded text2img generation with caption and bounding box condition inputs, and the grounding ability generalizes well to novel spatial configurations and concepts. GLIGEN's zero-shot performance on COCO and LVIS outperforms that of existing supervised layout-to-image baselines by a large margin.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries on Latent Diffusion Models
  • 4 Open-set Grounded Image Generation
  • 4.1 Grounding Instruction Input
  • 4.2 Continual Learning for Grounded Generation
  • 5 Experiments
  • 5.1 Closed-set Grounded Text2Img Generation
  • 5.2 Open-set Grounded Text2Img Generation
  • 5.3 Beyond Text Modality Grounding
  • 5.4 Scheduled Sampling
  • 6 Conclusion
  • References
  • A Implementation and training details
  • B Ablation Study
  • C Grounded inpainting
  • C.1 Text Grounded Inpainting
  • C.2 Image Grounded Inpainting
  • D Study for Keypoints Grounding
  • E Additional quantitative results
  • F Analysis on Gligen
  • G More qualitative results

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. ArXiv, abs/2204.14198, 2022.
  2. 2.Agrim Gupta, Piotr Dollar, and Ross B. Girshick. Lvis: A dataset for large vocabulary instance segmentation. CVPR, pages 5351–5359, 2019.
  3. 3.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross B. Girshick. Mask r-cnn. 2017 IEEE International Conference on Computer Vision (ICCV), pages 2980–2988, 2017.
  4. 4.Jonathan Ho. Classifier-free diffusion guidance. ArXiv, abs/2207.12598, 2022.
  5. 5.Manuel Jahn, Robin Rombach, and Bjorn Ommer. High-resolution complex scene synthesis with transformers. ArXiv, abs/2105.06458, 2021.
  6. 6.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2015.
  7. 7.Z. Li, Jingyu Wu, Immanuel Koh, Yongchuan Tang, and Lingyun Sun. Image synthesis from layout with locality-aware mask adaption. ICCV, pages 13799–13808, 2021.
  8. 8.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  9. 9.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  10. 10.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  11. 11.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022.
  12. 12.Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. CVPR, pages 10674–10685, 2022.
  13. 13.O. Ronneberger, P.Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI), volume 9351 of LNCS, pages 234–241. Springer, 2015. (available on arXiv:1505.04597 [cs.CV]).
  14. 14.Wei Sun and Tianfu Wu. Learning layout and style reconfigurable gans for controllable image synthesis. TPAMI, 44:5070–5087, 2022.
  15. 15.Tristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm, and Shikhar Sharma. Object-centric image generation from layouts. ArXiv, abs/2003.07449, 2021.
  16. 16.Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8798–8807, 2018.
  17. 17.Zuopeng Yang, Daqing Liu, Chaoyue Wang, J. Yang, and Dacheng Tao. Modeling image composition for complex scene generation. CVPR, pages 7754–7763, 2022.

Citation

MLA
Li, Y., et al. “GLIGEN: Open-Set Grounded Text-to-Image Generation”. arXiv, 2023, http://arxiv.org/abs/2301.07093v2.
APA
Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., & Lee, Y. J. (2023). GLIGEN: Open-Set Grounded Text-to-Image Generation. arXiv. http://arxiv.org/abs/2301.07093v2
Chicago
Li, Y., H. Liu, Q. Wu, et al. 2023. “GLIGEN: Open-Set Grounded Text-to-Image Generation”. arXiv. http://arxiv.org/abs/2301.07093v2.
Harvard
Li, Y. et al. (2023) “GLIGEN: Open-Set Grounded Text-to-Image Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.07093v2.
Vancouver
1. Li Y, Liu H, Wu Q, Mu F, Yang J, Gao J, Li C, Lee YJ (2023) GLIGEN: Open-Set Grounded Text-to-Image Generation. arXiv

BibTeX

@article{li2023gligen,
  title = {GLIGEN: Open-Set Grounded Text-to-Image Generation},
  author = {Li, Yuheng and Liu, Haotian and Wu, Qingyang and Mu, Fangzhou and Yang, Jianwei and Gao, Jianfeng and Li, Chunyuan and Lee, Yong Jae},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.07093v2},
  eprint = {2301.07093}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE