Kosmos-G: Generating Images in Context with Multimodal Large Language Models
Xichen PanLi DongShaohan HuangZhiliang PengWenhu ChenFuru Wei
Presents Kosmos-G, a framework that aligns multimodal large language models with CLIP to achieve zero-shot subject-driven image generation from interleaved text and image inputs without requiring test-time tuning or image decoder modifications.
Subject-driven image generation allows systems to create customized visuals based on reference images. However, prevailing methods require slow, computationally expensive fine-tuning at test time for each new subject and struggle to process inputs that mix multiple reference images with text descriptions. This creates a significant operational bottleneck for organizations seeking scalable, real-time image customization and fine-grained creative control.
The article demonstrates KOSMOS-G, an AI model that achieves zero-shot subject-driven image generation using interleaved text and multi-image inputs. The primary objective is to treat images as a "foreign language" within a unified generation pipeline, allowing the model to incorporate multiple novel visual concepts into new scenes without per-subject fine-tuning.
To achieve this, the authors implemented a three-stage "align before instruct" training approach. First, a 1.6-billion-parameter multimodal language model was pre-trained to perceive arbitrary sequences of text and images. Second, an alignment module was trained on text data to map the multimodal model's output space directly to the input space of a standard, frozen Stable Diffusion image generation engine using CLIP supervision. Third, the system underwent compositional instruction tuning on roughly 200 million curated examples using score distillation, training the model to faithfully reproduce segmented visual entities while leaving the underlying image generation network untouched.
Quantitative and qualitative evaluations reveal several key findings. First, KOSMOS-G achieves leading zero-shot subject fidelity on the standard DreamBench benchmark (0.694 DINO and 0.847 CLIP-I scores), outperforming tuning-free alternatives and matching or exceeding computationally heavy fine-tuning methods like DreamBooth using only a single reference image. Second, on standard text-to-image benchmarks, the model achieved a 10.99 Fréchet Inception Distance score, surpassing competing vision-language-to-image models. Third, ablation analyses showed that direct end-to-end training without the dedicated alignment network fails, confirming that explicit space alignment is essential for image quality. Finally, the model successfully demonstrated zero-shot generation across complex multi-entity prompts containing three to four distinct visual subjects.
These findings indicate that organizations can achieve highly customized, multi-subject image synthesis at zero-shot inference speeds, drastically reducing computing overhead and latency compared to per-subject fine-tuning pipelines. Because KOSMOS-G leaves the core diffusion model unchanged, it serves as a direct, drop-in replacement for standard text encoders. This ensures full plug-and-play compatibility with existing structural control tools like ControlNet and stylized visual adapters like LoRA.
Decision-makers should view this architecture as a viable pathway for high-throughput visual synthesis workflows that require multi-object personalization. Before commercial deployment, teams should conduct further development and piloting to address specific prompt sensitivities, such as formatting artifacts introduced by automated data captioning tools. Furthermore, rigorous content filtering and safety guardrails must be maintained, as the current release is strictly a research project with no immediate commercial availability.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). Introduces the foundation multimodal language model architecture and visual grounding techniques that Kosmos-G adapts as its core multimodal perception backbone.
- Paper: DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation, Nataniel Ruiz et al. (2023). Establishes the standard DreamBench evaluation framework and per-subject fine-tuning paradigm that Kosmos-G aims to replace with zero-shot generation.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). Provides fundamental insights into multi-subject concept composition in text-to-image diffusion models, motivating Kosmos-G's zero-shot multi-entity synthesis.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). Presents the concept of learning discrete embedding representations for visual subjects in text encoders, which informs Kosmos-G's strategy of mapping visual subjects into diffusion input space.
- Paper: IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models, Hu Ye et al. (2023). Demonstrates decoupled cross-attention and visual prompt adaptation for frozen diffusion models, establishing key principles for plug-and-play image conditioning.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). Introduces structural adapters that preserve a frozen base diffusion model, which directly relates to Kosmos-G maintaining modular drop-in compatibility with downstream tools.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). Formulates the Score Distillation Sampling loss technique that Kosmos-G employs during its compositional instruction tuning stage.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Builds upon unified multimodal understanding and generation paradigms by integrating discrete diffusion and autoregressive modelling inside a single unified transformer architecture.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). Extends multimodal in-context generation to foundation-scale autoregressive models capable of unified multi-task visual learning and generation.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). Generalizes multi-image conditioning and subject customization into a unified diffusion transformer framework trained on discontinuous video dynamics.
- Paper: SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models, Yuzhou Huang et al. (2024). Applies multimodal large language model perception to instruction-based image editing requiring complex contextual reasoning and bidirectional feature interaction.
- Paper: Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models, Chang Liu et al. (2024). Applies multimodal context conditioning to multi-frame visual storytelling, ensuring sequential character and narrative consistency across open-ended generations.
- Paper: VideoBooth: Diffusion-based Video Generation with Image Prompts, Yuming Jiang et al. (2024). Extends tuning-free subject customization concepts from single-image synthesis to temporally consistent diffusion-based video generation.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). Leverages multimodal LLMs to systematically assess semantic consistency and perceptual quality across conditional and subject-driven image synthesis tasks.
