Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Xichen PanLi DongShaohan HuangZhiliang PengWenhu ChenFuru Wei

article2024ICLR114 citations

Presents Kosmos-G, a framework that aligns multimodal large language models with CLIP to achieve zero-shot subject-driven image generation from interleaved text and image inputs without requiring test-time tuning or image decoder modifications.

Listen

Subject-driven image generation allows systems to create customized visuals based on reference images. However, prevailing methods require slow, computationally expensive fine-tuning at test time for each new subject and struggle to process inputs that mix multiple reference images with text descriptions. This creates a significant operational bottleneck for organizations seeking scalable, real-time image customization and fine-grained creative control.

The article demonstrates KOSMOS-G, an AI model that achieves zero-shot subject-driven image generation using interleaved text and multi-image inputs. The primary objective is to treat images as a "foreign language" within a unified generation pipeline, allowing the model to incorporate multiple novel visual concepts into new scenes without per-subject fine-tuning.

To achieve this, the authors implemented a three-stage "align before instruct" training approach. First, a 1.6-billion-parameter multimodal language model was pre-trained to perceive arbitrary sequences of text and images. Second, an alignment module was trained on text data to map the multimodal model's output space directly to the input space of a standard, frozen Stable Diffusion image generation engine using CLIP supervision. Third, the system underwent compositional instruction tuning on roughly 200 million curated examples using score distillation, training the model to faithfully reproduce segmented visual entities while leaving the underlying image generation network untouched.

Quantitative and qualitative evaluations reveal several key findings. First, KOSMOS-G achieves leading zero-shot subject fidelity on the standard DreamBench benchmark (0.694 DINO and 0.847 CLIP-I scores), outperforming tuning-free alternatives and matching or exceeding computationally heavy fine-tuning methods like DreamBooth using only a single reference image. Second, on standard text-to-image benchmarks, the model achieved a 10.99 Fréchet Inception Distance score, surpassing competing vision-language-to-image models. Third, ablation analyses showed that direct end-to-end training without the dedicated alignment network fails, confirming that explicit space alignment is essential for image quality. Finally, the model successfully demonstrated zero-shot generation across complex multi-entity prompts containing three to four distinct visual subjects.

These findings indicate that organizations can achieve highly customized, multi-subject image synthesis at zero-shot inference speeds, drastically reducing computing overhead and latency compared to per-subject fine-tuning pipelines. Because KOSMOS-G leaves the core diffusion model unchanged, it serves as a direct, drop-in replacement for standard text encoders. This ensures full plug-and-play compatibility with existing structural control tools like ControlNet and stylized visual adapters like LoRA.

Decision-makers should view this architecture as a viable pathway for high-throughput visual synthesis workflows that require multi-object personalization. Before commercial deployment, teams should conduct further development and piloting to address specific prompt sensitivities, such as formatting artifacts introduced by automated data captioning tools. Furthermore, rigorous content filtering and safety guardrails must be maintained, as the current release is strictly a research project with no immediate commercial availability.

arXiv: 2310.02992
Cover for Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Abstract

Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." The code can be found at this https URL

Table of Contents

  • 1 Introduction
  • 2 Kosmos-G: Image as a Foreign Language in Image Generation
  • 2.1 Multimodal Language Modeling
  • 2.2 Image Decoder Aligning
  • 2.3 Instruction Tuning
  • 3 Model Training
  • 3.1 Multimodal Training Data
  • 3.2 Training Setup
  • 4 Evaluation
  • 4.1 Main Qualitative Results
  • 4.2 Quantitative Results
  • 4.3 Ablation Studies
  • 4.4 Applications
  • 5 Conclusion
  • References
  • A Detail Training Setup
  • B Images and Prompts for DreamBench Evaluation
  • C Additional Details about Score Distillation Instruction Tuning
  • D Additional Examples

Knowls

  1. Knowl 1 — KOSMOS-G Architecture and Three-Stage Training Framework

    model/method

    KOSMOS-G is a multimodal generative framework capable of zero-shot subject-driven image generation from interleaved multi-image and text prompts. It connects a Multimodal Large Language Model (MLLM) to a frozen diffusion image decoder (the U-Net from Stable Diffusion v1.5) via an intermediate bridging module called AlignerNet, trained through a three-stage "align before instruct" pipeline:

    1. Multimodal Language Modeling: An MLLM is pre-trained from scratch using next-token causal language modeling on monomodal text corpora, cross-modal image-caption pairs, and interleaved multimodal sequences. The language backbone contains 24 MAGNETO Transformer layers with a hidden dimension of 2,048, an intermediate FFN dimension of 8,192, and 32 attention heads (totaling approximately 1.6B parameters). Visual inputs are processed at 224×224224 \times 224 resolution by a CLIP ViT-L/14 vision encoder (1,024 feature dimensions) followed by an attentive Resampler pooling module.

    2. Image Decoder Aligning: With the MLLM and the diffusion U-Net both frozen, AlignerNet is trained exclusively on text captions using CLIP text encoder supervision. Text acts as an anchoring modality, mapping the output space of KOSMOS-G into the cross-attention conditioning space of the diffusion U-Net.

    3. Compositional Instruction Tuning: The MLLM and AlignerNet are jointly fine-tuned on curated multi-entity segmentation and image-editing data using score distillation gradients backpropagated from the frozen diffusion U-Net, enabling faithful multi-subject image generation without altering the diffusion weights.

  2. Knowl 2 — AlignerNet Architecture and Text-Anchored Output Space Alignment

    model/method

    AlignerNet bridges the variable-length representation space S\mathcal{S} produced by KOSMOS-G to the fixed target conditioning space T\mathcal{T} of the pre-trained CLIP text encoder used by Stable Diffusion. It consists of an encoder M\mathcal{M} and a decoder N\mathcal{N}, each comprising a 12-layer Transformer encoder and a 12-layer Transformer decoder with input dimension d=768d = 768 and hidden dimension 3,0723,072 (totaling roughly 225M parameters).

    To bridge variable-length input sequences to target lengths, the cross-attention layers of M\mathcal{M} and N\mathcal{N} employ learned latent query matrices QM∈Rlt×dQ_{\mathcal{M}} \in \mathbb{R}^{l_t \times d} and QN∈Rls×dQ_{\mathcal{N}} \in \mathbb{R}^{l_s \times d}, where lsl_s is the source token length and ltl_t is the target CLIP token length.

    Given source embedding s∈Rls×dss \in \mathbb{R}^{l_s \times d_s} from KOSMOS-G and target embedding t∈Rlt×dtt \in \mathbb{R}^{l_t \times d_t} from the CLIP text encoder for a caption, AlignerNet is trained on text pairs with two loss objectives:

    Lmse=Es∼S,t∼T[∥t−M(s)∥22]\mathcal{L}_{\text{mse}} = \mathbb{E}_{s \sim \mathcal{S}, t \sim \mathcal{T}} \left[ \| t - \mathcal{M}(s) \|_2^2 \right]

    Lrec=Es∼S[∥s−N(M(s))∥22]\mathcal{L}_{\text{rec}} = \mathbb{E}_{s \sim \mathcal{S}} \left[ \| s - \mathcal{N}(\mathcal{M}(s)) \|_2^2 \right]

    Because language serves as the anchor during this stage, the vision representations that were aligned with language during MLLM pre-training are naturally aligned into the diffusion conditioning space.

  3. Knowl 3 — Score Distillation Instruction Tuning Formulation for Subject Conditioning

    equation

    To train KOSMOS-G to preserve fine concept-level entity details from reference images without fine-tuning the diffusion image decoder, the model utilizes score distillation from the pre-trained frozen Stable Diffusion U-Net ϵθ\epsilon_\theta.

    Let ϕ\phi represent KOSMOS-G (the MLLM combined with AlignerNet), which maps an interleaved multimodal input prompt xx to a conditioning representation C=ϕ(x)C = \phi(x). The diffusion loss over latent state z0=E(xtarget)z_0 = \mathcal{E}(x_{\text{target}}) encoded by VAE encoder E\mathcal{E} is:

    Ldiff(ϕ)=Ez0,ϵ∼N(0,I),t[w(t)∥ϵθ(zt;C,t)−ϵ∥22]\mathcal{L}_{\text{diff}}(\phi) = \mathbb{E}_{z_0, \epsilon \sim \mathcal{N}(0, I), t} \left[ w(t) \| \epsilon_\theta(z_t; C, t) - \epsilon \|_2^2 \right]

    where t∈[1,T]t \in [1, T] is the diffusion time step, zt=αˉtz0+1−αˉtϵz_t = \sqrt{\bar{\alpha}_t} z_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon is the noisy latent with schedule αˉt=∏i=1t(1−βi)\bar{\alpha}_t = \prod_{i=1}^t (1 - \beta_i), ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is added Gaussian noise, and w(t)w(t) is a time-dependent weighting coefficient.

    Omitting the computationally expensive U-Net Jacobian ∂ϵθ(zt;C,t)∂C\frac{\partial \epsilon_\theta(z_t; C, t)}{\partial C} yields the Score Distillation Sampling (SDS) gradient applied to update ϕ\phi:

    ∇ϕLSDS(ϕ)=Ez0,ϵ∼N(0,I),t[w(t)(ϵθ(zt;C,t)−ϵ)∂C∂ϕ]\nabla_\phi \mathcal{L}_{\text{SDS}}(\phi) = \mathbb{E}_{z_0, \epsilon \sim \mathcal{N}(0, I), t} \left[ w(t) (\epsilon_\theta(z_t; C, t) - \epsilon) \frac{\partial C}{\partial \phi} \right]

    Optimizing this objective minimizes the Kullback-Leibler divergence between the forward diffusion distribution and the score function parameterized by ϕ\phi:

    min⁡ϕLdiff(ϕ)=Ez0,t,C[DKL(q(zt−1∣zt,z0)∥pθ(zt−1∣zt;C))]\min_\phi \mathcal{L}_{\text{diff}}(\phi) = \mathbb{E}_{z_0, t, C} \left[ D_{\text{KL}}\left(q(z_{t-1} \mid z_t, z_0) \parallel p_\theta(z_{t-1} \mid z_t; C)\right) \right]

  4. Knowl 4 — Automated Data Curation Algorithm for Compositional Instruction Tuning

    algorithm

    To train KOSMOS-G to follow multi-entity multimodal prompts, an automated segmentation and formatting pipeline constructs training samples from roughly 9 million images from the Open Images V7 dataset.

    Input: Unlabeled image dataset D_images
    Output: Compositional multimodal instruction dataset D_instruct
    D_instruct = []
    for each image I in D_images do
        caption = BLIP_2_OPT_6_7B(I)
        entities = MPT_7B_Instruct_ExtractEntities(caption)
        interleaved_tokens = []
        for each entity_text in entities do
            mask = CLIPSeg(I, text_prompt=entity_text)
            crop = ApplySegmentation(I, mask)
            if UniformRandom(0, 1) < 0.5 then
                crop = KeepBackground(crop, I)
            end if
            if UniformRandom(0, 1) < 0.5 then
                entity_text = "" # entity text dropout
            end if
            crop_emb = Resampler(CLIP_ViT(crop))
            interleaved_tokens.append(entity_text, "<image>", crop_emb, "</image>")
        end for
        prompt = AssembleSequence("<s>", caption_with_interleaved_tokens, "</s>")
        D_instruct.append((prompt, I))
    end for
    return D_instruct

    During instruction tuning, the resulting dataset is blended with InstructPix2Pix image editing data and standard image-caption pairs in a 2:2:1 ratio.

  5. Knowl 5 — Subject-Driven Zero-Shot Generation Performance on DreamBench

    data/table

    KOSMOS-G was evaluated on DreamBench across 30 subjects with 25 prompt templates (750 unique prompts, 4 samples each, totaling 3,000 generated images) using single-image reference inputs. Subject fidelity was measured using DINO and CLIP-I embedding similarities; text alignment was measured via CLIP-T score. Inference used 100 DPM-Solver steps with a classifier-free guidance scale of 7.5.

    Methods DINO ↑\uparrow CLIP-I ↑\uparrow CLIP-T ↑\uparrow
    Real Images (Oracle) 0.774 0.885 -
    Fine-Tuning Methods
    Textual Inversion 0.569 0.780 0.255
    DreamBooth 0.668 0.803 0.305
    BLIP-Diffusion 0.670 0.805 0.302
    Test-Time Tuning-Free Methods
    Re-Imagen* 0.600 0.740 0.270
    SuTI 0.741 0.819 0.304
    BLIP-Diffusion* 0.594 0.779 0.300
    KOSMOS-G* (single image input) 0.694 0.847 0.287
    • denotes zero-shot methods requiring no test-time tuning. KOSMOS-G achieves a CLIP-I score of 0.847, outperforming all test-time fine-tuning methods (Textual Inversion, DreamBooth, BLIP-Diffusion) and tuning-free baselines while achieving comparable subject consistency to SuTI without requiring apprenticeship training.
  6. Knowl 6 — Zero-Shot Text-to-Image Generation Quality on MS-COCO

    data/table

    The zero-shot text-to-image synthesis quality of KOSMOS-G was evaluated on 30,000 randomly selected captions from the MS-COCO 2014 validation set using 250 DDIM inference steps and a classifier-free guidance scale of 3.0, reporting Fréchet Inception Distance (FID).

    Methods FID ↓\downarrow
    T2I Models
    GLIDE 12.24
    Make-A-Scene 11.84
    DALL-E 2 10.39
    SD v1.5 9.34
    Imagen-3.4B 7.27
    CLIP-Aligned VL2I Models
    GILL-8B 12.20
    Emu-14B 11.66
    KOSMOS-G-1.9B 10.99

    KOSMOS-G-1.9B achieves an FID of 10.99, outperforming larger multimodal CLIP-aligned models such as GILL-8B (12.20) and Emu-14B (11.66) while retaining the original text-to-image generation fidelity of the underlying diffusion backbone.

  7. Knowl 7 — Ablation Study on Decoder Aligning Architecture and Training Strategy

    data/table

    Ablation experiments evaluate the contribution of AlignerNet and the intermediate decoder alignment stage on MS-COCO zero-shot text-to-image FID:

    Methods FID ↓\downarrow
    SD v1.5 Baseline 9.34
    End-to-End w/o AlignerNet Failed
    End-to-End w/ AlignerNet 11.30
    12-Layers Decoder-only Failed
    12-Layers AlignerNet 9.89
    24-Layers AlignerNet 9.55

    Attempting direct end-to-end training without AlignerNet or using an asymmetric decoder-only bridge fails completely to produce coherent images. Using an encoder-decoder AlignerNet with text supervision aligns the representation spaces effectively, with a 24-layer AlignerNet reaching an FID of 9.55, near the upper bound of the frozen SD v1.5 baseline (9.34).

  8. Knowl 8 — Plug-and-Play Compatibility with Diffusion U-Net Control and Personalization Extensions

    model/method

    Because KOSMOS-G aligns its output directly to the cross-attention input space of the diffusion model without modifying any parameters of the Stable Diffusion U-Net or VAE, it acts as a drop-in replacement for the original CLIP text encoder. This enables direct interoperability with external U-Net modules without fine-tuning:

    1. Spatial Structure Control: KOSMOS-G can be combined directly with ControlNet (such as Canny edge conditioning) to specify exact spatial geometry while using multimodal interleaved prompts for subject identity and styling.
    2. Personalized and Stylized Decoders: KOSMOS-G interfaces with pre-trained Low-Rank Adaptation (LoRA) weights injected into the U-Net, applying distinct artistic styles or subject checkpoints alongside in-context multi-image prompts.

Coverage note — None was omitted; all core contributions, algorithmic procedures, mathematical formulations, experimental benchmarks, ablation results, and application demonstrations are covered.

References

  1. 1.Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel Cohen-Or, and Dani Lischinski. Break-a-scene: Extracting multiple concepts from a single image. arXiv preprint arXiv:2305.16311, 2023.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems, 2022.
  3. 3.Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. CM3: A causal masked multimodal model of the Internet. ArXiv preprint, abs/2201.07520, 2022.
  4. 4.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023.
  5. 5.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset, 2022.
  6. 6.Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Rui, Xuhui Jia, Ming-Wei Chang, and William W Cohen. Subject-driven text-to-image generation via apprenticeship learning. ArXiv preprint, abs/2304.00186, 2023.
  7. 7.Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W Cohen. Re-imagen: Retrieval-augmented text-to-image generator. ArXiv preprint, abs/2209.14491, 2022.
  8. 8.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 3558–3568. Computer Vision Foundation / IEEE, 2021.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  10. 10.Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, et al. Dreamllm: Synergistic multimodal comprehension and creation. ArXiv preprint, abs/2309.11499, 2023.
  11. 11.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. ArXiv preprint, abs/2208.01618, 2022.
  12. 12.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. ArXiv preprint, abs/2203.13131, 2022.
  13. 13.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al. Language is not all you need: Aligning perception with language models. ArXiv preprint, abs/2302.14045, 2023.
  14. 14.Shaozhe Hao, Kai Han, Shihao Zhao, and Kwan-Yee K Wong. Vico: Detail-preserving visual condition for personalized text-to-image generation. arXiv preprint arXiv:2306.00971, 2023.
  15. 15.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  16. 16.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. ArXiv preprint, abs/2207.12598, 2022.
  17. 17.Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei. Language models are general-purpose interfaces. ArXiv preprint, abs/2206.06336, 2022.
  18. 18.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022.
  19. 19.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. Generating images with multimodal language models. ArXiv preprint, abs/2305.17216, 2023.
  20. 20.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium, 2018. Association for Computational Linguistics.
  21. 21.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020.
  22. 22.Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In Yoshua Bengio and Yann LeCun, editors, 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014.
  23. 23.Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1931–1941, 2023.
  24. 24.Timo Lüddecke and Alexander Ecker. Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7086–7096, 2022.
  25. 25.Dongxu Li, Junnan Li, and Steven CH Hoi. Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. ArXiv preprint, abs/2305.14720, 2023.
  26. 26.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. ArXiv preprint, abs/2301.12597, 2023.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized bert pretraining approach. ArXiv preprint, abs/1907.11692, 2019.
  29. 29.Calvin Luo. Understanding diffusion models: A unified perspective. arXiv preprint arXiv:2208.11970, 2022.
  30. 30.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Advances in Neural Information Processing Systems, 35:5775–5787, 2022.
  31. 31.Shuming Ma, Hongyu Wang, Shaohan Huang, Wenhui Wang, Zewen Chi, Li Dong, Alon Benhaim, Barun Patra, Vishrav Chaudhary, Xia Song, and Furu Wei. TorchScale: Transformers at scale. ArXiv preprint, abs/2211.13184, 2022.
  32. 32.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 16784–16804. PMLR, 2022.
  33. 33.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. ArXiv preprint, abs/2209.14988, 2022.
  34. 34.Can Qin, Ning Yu, Chen Xing, Shu Zhang, Zeyuan Chen, Stefano Ermon, Yun Fu, Caiming Xiong, and Ran Xu. Gluegen: Plug and play multi-modal encoders for x-to-image generation. ArXiv preprint, abs/2303.10056, 2023.
  35. 35.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022.
  36. 36.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. ArXiv preprint, abs/2204.06125, 2022.
  37. 37.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021.
  39. 39.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. ArXiv preprint, abs/2208.12242, 2022.
  40. 40.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. ArXiv preprint, abs/2210.08402, 2022.
  41. 41.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. ArXiv preprint, abs/2205.11487, 2022.
  42. 42.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, Melbourne, Australia, 2018. Association for Computational Linguistics.
  43. 43.James Seale Smith, Yen-Chang Hsu, Lingyu Zhang, Ting Hua, Zsolt Kira, Yilin Shen, and Hongxia Jin. Continual diffusion: Continual customization of text-to-image diffusion with c-lora. arXiv preprint arXiv:2304.06027, 2023.
  44. 44.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  45. 45.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. ArXiv preprint, abs/2111.02114, 2021.
  46. 46.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In Francis R. Bach and David M. Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pages 2256–2265. JMLR.org, 2015.
  47. 47.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pre-training in multimodality. ArXiv preprint, abs/2307.05222, 2023.
  48. 48.MN Team et al. Introducing MPT-7B: A new standard for open-source, commercially usable llms, 2023.
  49. 49.Yoad Tewel, Rinon Gal, Gal Chechik, and Yuval Atzmon. Key-locked rank one editing for text-to-image personalization. arXiv preprint arXiv:2305.01644, 2023.
  50. 50.Hongyu Wang, Shuming Ma, Shaohan Huang, Li Dong, Wenhui Wang, Zhiliang Peng, Yu Wu, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Xia Song, and Furu Wei. Foundation transformers. ArXiv preprint, abs/2210.06423, 2022.
  51. 51.Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. arXiv preprint arXiv:2302.13848, 2023.
  52. 52.Guangxuan Xiao, Tianwei Yin, William T Freeman, Frédo Durand, and Song Han. Fastcomposer: Tuning-free multi-subject image generation with localized attention. arXiv preprint arXiv:2305.10431, 2023.
  53. 53.Lvmin Zhang and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023.

Citation

MLA
Pan, X., et al. “Kosmos-G: Generating Images in Context with Multimodal Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2310.02992v3.
APA
Pan, X., Dong, L., Huang, S., Peng, Z., Chen, W., & Wei, F. (2023). Kosmos-G: Generating Images in Context with Multimodal Large Language Models. arXiv. http://arxiv.org/abs/2310.02992v3
Chicago
Pan, X., L. Dong, S. Huang, Z. Peng, W. Chen, and F. Wei. 2023. “Kosmos-G: Generating Images in Context with Multimodal Large Language Models”. arXiv. http://arxiv.org/abs/2310.02992v3.
Harvard
Pan, X. et al. (2023) “Kosmos-G: Generating Images in Context with Multimodal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.02992v3.
Vancouver
1. Pan X, Dong L, Huang S, Peng Z, Chen W, Wei F (2023) Kosmos-G: Generating Images in Context with Multimodal Large Language Models. arXiv

BibTeX

@article{pan2023kosmos,
  title = {Kosmos-G: Generating Images in Context with Multimodal Large Language Models},
  author = {Pan, Xichen and Dong, Li and Huang, Shaohan and Peng, Zhiliang and Chen, Wenhu and Wei, Furu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.02992v3},
  eprint = {2310.02992}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors