GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng LiHaotian LiuQingyang WuFangzhou MuJianwei YangJianfeng GaoChunyuan LiYong Jae Lee
Proposes GLIGEN, a framework that injects spatial grounding inputs into frozen pre-trained diffusion models via gated layers, enabling precise layout-controlled image generation across open-world concepts without retraining the base model.
Existing text-to-image diffusion models generate impressive visual content from open-ended prompts, but they frequently struggle with precise spatial controllability, object placement, and editing accuracy. In standard text-to-image workflows, users cannot reliably specify exact object locations, poses, or boundaries, which limits their utility in production environments that require fine-grained scene composition.
The article evaluates GLIGEN (Grounded-Language-to-Image Generation), an approach designed to endow text-to-image diffusion models with open-set grounded generation capabilities using bounding boxes, human keypoints, reference images, and spatial maps.
The authors integrated grounded condition tokens into pretrained diffusion backbones using newly added gated self-attention layers while keeping the core generative weights intact. The evaluation spanned multiple benchmark datasets (COCO, LVIS, GoldG, Objects365, CC3M, and SBU) and tested tasks including text-grounded inpainting, image-grounded editing, keypoint-guided human generation, and layout-to-image synthesis.
The experiments show that incorporating spatial grounding substantially improves alignment and image fidelity across tasks. In text-grounded inpainting, the method achieved higher precision across all object sizes compared to baseline latent diffusion, maintaining an average precision of 25.6% on large objects where the baseline dropped to 14.6%. For human keypoint conditioning on COCO, the approach reached an average precision of 31.8% and a Fréchet Inception Distance of 31.02, dramatically outperforming the dedicated translation baseline pix2pixHD (15.8% AP and 142.4 FID). Pretraining across large-scale datasets and finetuning on LVIS yielded an average precision of 14.9% and an FID of 6.25, significantly surpassing supervised layout models like LAMA. Furthermore, gated self-attention proved superior to cross-attention mechanisms by allowing necessary visual token interaction.
These results demonstrate that spatial grounding can be added efficiently to pretrained models without retraining foundational weights from scratch. This reduces computational training costs and operational risk while providing reliable layout control for commercial creative tools, precision editing, and synthetic data pipelines.
Organizations developing controllable generative media should adopt gated self-attention architectures and leverage mixed caption and detection datasets during pretraining. When deploying spatial conditioning, teams should pair coarse bounding boxes for category-agnostic layout control and reserve keypoint conditioning for specific humanoid poses.
A key limitation identified is that keypoint grounding does not generalize across non-humanoid object categories, as anatomical part associations are less transferable than generic bounding boxes. Confidence in the reported image quality and layout correspondence remains high across standard bounding box and human pose benchmarks.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). ControlNet establishes the key design precedent of adding visual conditioning to a frozen text-to-image diffusion backbone, clarifying the control strategy GLIGEN adapts for grounded inputs.
- Paper: Boximator: Generating Rich and Controllable Motions for Video Synthesis, Jiawei Wang et al. (2024). Boximator carries spatial box conditioning from still-image generation into video, adding trajectories and temporal consistency to the grounded-control problem GLIGEN addresses.
