Generative Multimodal Models are In-Context Learners

Quan SunYufeng CuiXiaosong ZhangFan ZhangQiying YuYueze WangYongming RaoJingjing LiuTiejun HuangXinlong Wang

article2024CVPR541 citations

Introduces Emu2, a 37-billion-parameter generative multimodal model trained with a unified autoregressive objective that achieves state-of-the-art in-context learning and controllable generation across diverse visual-language tasks.

Listen

Multimodal artificial intelligence systems often require customized architectures and dedicated supervised datasets for every new application, making deployment across diverse real-world tasks difficult and costly. By contrast, humans can learn new concepts or solve complex tasks on the fly using only a few demonstrations or basic instructions. While large language models have mastered this capability in text, multimodal models have lagged behind in flexibly handling mixed sequences of images, video, and text.

The article demonstrates that scaling up a generative multimodal model using a unified learning objective enables strong in-context learning and generalization across both visual understanding and visual content generation. Specifically, it presents and evaluates Emu2, a 37-billion-parameter model designed to operate as a general-purpose multimodal foundation system.

The authors constructed Emu2 by integrating a visual encoder, a multimodal language model backbone, and a visual diffusion decoder, training the entire system with an autoregressive "predict-the-next-element" objective across large-scale text, paired image-video data, and interleaved sequences. The framework was evaluated across two main regimes: few-shot in-context learning with increasing numbers of visual-text examples, and task-specific instruction tuning for dialogue (Emu2-Chat) and controllable image generation (Emu2-Gen). Evaluations encompassed standard visual question answering, referring expression comprehension, complex multimodal reasoning benchmarks, and subject-driven image generation.

Key findings show that Emu2 sets new performance benchmarks while requiring fewer parameters than leading competitors. In few-shot evaluations, Emu2's accuracy consistently improved as context examples grew, outperforming larger 80-billion-parameter models like Flamingo-80B and IDEFICS-80B on benchmarks such as VQAv2 (68.8% vs. 66.8% at 16 shots) and TextVQA (50.3% vs. 37.6%). When instruction-tuned, Emu2-Chat attained state-of-the-art results on comprehensive evaluation suites, scoring 48.5 on MM-Vet and 703.8 on TouchStone, while also leading generalist models in referring expression comprehension on RefCOCO datasets. Additionally, in controllable visual generation, Emu2-Gen achieved industry-leading image and text alignment scores on MS-COCO zero-shot generation and produced superior subject fidelity on DreamBench compared to dedicated tuning-free alternatives.

These results demonstrate that a single, unified generative model can effectively bridge perception and generation, eliminating the need to maintain separate, narrow systems for visual question answering, grounding, and content creation. For organizations, this approach offers substantial efficiency gains and operational simplicity by providing a single interface for diverse multimodal workflows. However, zero-shot performance without prior context examples remains a relative weakness for the base model, emphasizing that contextual prompting or instruction tuning is necessary to unlock optimal performance.

Decision-makers and research teams should consider adopting unified generative architectures as base models for integrated visual and linguistic workflows, while establishing clear instruction-tuning pipelines for specific production tasks. Further work should focus on addressing societal misuse risks, mitigating potential biases inherent in web-scale pretraining datasets, and optimizing inference efficiency before rolling out the system to latency-sensitive or safety-critical operational environments.

arXiv: 2312.13286
Cover for Generative Multimodal Models are In-Context Learners

Abstract

The human ability to easily solve multimodal tasks in context (i.e., with only a few demonstrations or simple instructions), is what current multimodal systems have largely struggled to imitate. In this work, we demonstrate that the task-agnostic in-context learning capabilities of large multimodal models can be significantly enhanced by effective scaling-up. We introduce Emu2, a generative multimodal model with 37 billion parameters, trained on large-scale multimodal sequences with a unified autoregressive objective. Emu2 exhibits strong multimodal in-context learning abilities, even emerging to solve tasks that require on-the-fly reasoning, such as visual prompting and object-grounded generation. The model sets a new record on multiple multimodal understanding tasks in few-shot settings. When instruction-tuned to follow specific instructions, Emu2 further achieves new state-of-the-art on challenging tasks such as question answering benchmarks for large multimodal models and open-ended subject-driven generation. These achievements demonstrate that Emu2 can serve as a base model and general-purpose interface for a wide range of multimodal tasks. Code and models are publicly available to facilitate future research.

Table of Contents

  • 1. Introduction
  • 2. Approach
  • 2.1. Model Architecture
  • 2.2. Pretraining
  • 2.2.1 Data
  • 2.2.2 Training
  • 2.3. Instruction Tuning
  • 2.2.3 Visual Decoding
  • 2.3.1 Instruction-Following Chat
  • 2.3.2 Controllable Visual Generation
  • 3. Evaluation
  • 3.1. Pretrained Base Model
  • 3.2. Instruction-Following Chat
  • 3.3. Controllable Visual Generation
  • 4. Related Work
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Emu2 Architecture

    model/method

    Emu2 is a 37-billion-parameter generative multimodal foundation model designed to process and generate interleaved sequences of text, images, and video within a single autoregressive framework. The architecture consists of three main components:

    1. Visual Encoder: An EVA-02-CLIP-E-plus model that maps each input image into continuous visual feature maps.
    2. Multimodal Modeling Backbone: A 33-billion-parameter Transformer decoder initialized from LLaMA-33B. To connect the Visual Encoder to the LLM backbone, Emu2 applies spatial 2D mean pooling to downsample the visual feature maps into an 8×88 \times 8 grid (N=64N = 64 visual tokens per image), followed by a single linear projection layer mapping the visual tokens into the embedding dimension of the LLM. This eliminates complex intermediate connectors such as C-Formers.
    3. Visual Decoder: A diffusion-based module initialized from SDXL (Stable Diffusion XL base) for images and Stable Diffusion 2.1 for video. During inference, the visual embeddings output by the autoregressive backbone are passed directly to the Visual Decoder, which reconstructs high-resolution images (1024×10241024 \times 1024) or video sequences on the fly.
  2. Knowl 2 — Emu2 Pretraining Protocol and Objectives

    model/method

    Emu2 is pretrained using a unified predict-the-next-multimodal-element autoregressive objective across interleaved multimodal sequences in two sequential stages:

    1. Stage 1 (Image/Video-Text Warm-up): The model is trained on 162 million image-text pairs (from LAION-2B and CapsFusion-120M) and 7 million video-text pairs (from WebVid-10M) with only standard cross-entropy captioning loss applied to textual tokens. Input images are resized to 224×224224 \times 224. The optimization uses AdamW (β1=0.9,β2=0.95,ϵ=1×10−6\beta_1 = 0.9, \beta_2 = 0.95, \epsilon = 1 \times 10^{-6}) with peak learning rates of 1×10−41 \times 10^{-4} for the linear projection layer, 3×10−53 \times 10^{-5} for Multimodal Modeling, and 5×10−55 \times 10^{-5} for the Visual Encoder. Training runs for 35,200 iterations with a global batch size of 6,144 for image-text pairs and 768 for video-text pairs, followed by a restart at 448×448448 \times 448 resolution for 4,000 additional iterations.
    2. Stage 2 (Unified Multimodal Pretraining): The Visual Encoder is frozen, and only the linear projection layer and Multimodal Modeling backbone are optimized using both text token classification loss (cross-entropy) and continuous visual embedding regression loss (mean squared error / L2L_2 regression on visual feature embeddings). Datasets include image-text pairs, video-text pairs, interleaved image-text data (Multimodal-C4), interleaved video-text data (YT-Storyboard-1B), grounded image-text pairs (GRIT-20M, CapsFusion-grounded-100M), and 3.8 billion language-only tokens from The Pile. All images are resized to 448×448448 \times 448. Training runs for 20,350 iterations at a maximum learning rate of 1×10−51 \times 10^{-5} with global batch sizes of 12,800 (image-text pairs), 6,400 (video-text pairs), 3,200 (interleaved data), and 800 (language-only data).
  3. Knowl 3 — Visual Decoder and Detokenization Architecture

    model/method

    Emu2 structures visual decoding as an off-the-shelf detokenizer trained independently of the autoregressive language model:

    • Image Detokenizer: Initialized from SDXL-base. The Visual Encoder (EVA-02-CLIP-E-plus) and SDXL VAE are kept frozen, while only the diffusion U-Net is trained under an autoencoding objective on LAION-COCO and LAION-Aesthetics. The cross-attention projection layers in the U-Net are resized to accept NN continuous visual embeddings from the Visual Encoder as conditional inputs. Input images are 448×448448 \times 448, and generated target images are 1024×10241024 \times 1024. Training employs classifier-free guidance by dropping visual embeddings with a 10% probability, using AdamW (β1=0.9,β2=0.999\beta_1 = 0.9, \beta_2 = 0.999, weight decay 0.01) with a peak learning rate of 1×10−41 \times 10^{-4} (2,000 warm-up steps, 6,000 linear decay steps) and a batch size of 2,048.
    • Video Detokenizer: Initialized from Stable Diffusion 2.1, adapting the 2D denoising U-Net into a pseudo-3D architecture by inserting a 1D temporal convolution after each 2D spatial convolution layer and extending spatial self-attention to spatio-temporal attention.
  4. Knowl 4 — Emu2-Chat Instruction Tuning

    model/method

    Emu2 is aligned for multimodal dialogue and instruction-following to create Emu2-Chat using the following setup:

    • Data Formatting: Training sequences are organized using special role tokens in the template: <Sys.Msg.> [USER]: <Instruction> [ASSISTANT]: <Answer>\text{<Sys.Msg.> [USER]: <Instruction> [ASSISTANT]: <Answer>} Cross-entropy loss is computed exclusively on tokens within the <Answer> section. Training data combines academic-task-oriented datasets (image captioning, VQA, science QA, multimodal classification, referring expressions) and conversational instruction datasets (GPT-assisted visual instructions, language instructions, clock reading, video chat), separated by system message definitions.
    • Visual Resolution and Tokenization: During instruction tuning, input static images (448×448448 \times 448) are mean-pooled into a 16×1616 \times 16 grid (256256 visual tokens per image) rather than the 8×88 \times 8 grid (6464 tokens) used in pretraining, providing higher spatial resolution. For video inputs, 8, 12, or 16 frames are uniformly sampled over time.
    • Optimization Hyperparameters: Global batch size of 768 for 8,000 steps, maximum sequence length 2,048 tokens. The learning rate linearly warms up to 1×10−51 \times 10^{-5} across the first 100 steps and decays to zero via a cosine schedule using AdamW (β1=0.9,β2=0.98,ϵ=1×10−6\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 1 \times 10^{-6}, gradient clipping norm 5.0).
  5. Knowl 5 — Emu2-Gen Controllable Visual Generation

    model/method

    Emu2 is adapted into Emu2-Gen for in-context controllable image generation (such as subject-driven generation, grounded text-to-image synthesis, image editing, and stylization) via unified multimodal autoregression:

    • Visual Coordinate and Conditioning Representation: Object locations are encoded not as textual coordinate numbers, but visually by rendering the bounding box of each object onto a black image. The input sequence represents entities, bounding box images, and crop images interleaved with text: <s>A photo of <p>man</p><coor>[BBox Image]</coor>[IMG][Man Crop][/IMG] sitting next to <p>dog</p><coor>[BBox Image]</coor>[IMG][Dog Crop][/IMG][IMG][Target Full Image][/IMG]</s>\text{<s>A photo of <p>man</p><coor>[BBox Image]</coor>[IMG][Man Crop][/IMG] sitting next to <p>dog</p><coor>[BBox Image]</coor>[IMG][Dog Crop][/IMG][IMG][Target Full Image][/IMG]</s>}
    • Training Strategy: The Visual Encoder is frozen. Continuous visual embedding regression loss is applied solely to the visual tokens of the final target image ([Target Full Image]). Entity text tokens and coordinate images are randomly dropped during training to improve robustness. Data augmentation on subject crops includes background replacement and random crops via Segment Anything (SAM) segmentation masks.
    • Optimization: Trained with a batch size of 4,096 for 3,000 steps (learning rate warming up to 5×10−55 \times 10^{-5} for 100 steps, cosine decaying to zero), followed by 900 fine-tuning steps on 500,000 high-quality text-image pairs (Unsplash, Midjourney-V5, DALL-E 3) with a batch size of 2,048.
  6. Knowl 6 — Few-Shot In-Context Multimodal Understanding Performance

    data/table

    Base Emu2 (37B parameters) demonstrates strong multimodal in-context learning capabilities across standard vision-language benchmarks, with performance scaling as the number of in-context demonstration shots increases.

    Model Shot VQAv2 OKVQA VizWiz TextVQA Hateful Memes
    Kosmos-1 (1.6B) 0 51.0 - 29.2 - -
    4 51.8 - 35.3 - -
    8 51.4 - 39.0 - -
    Flamingo (9B) 0∗0^* 51.8 44.7 28.8 31.8 57.0
    4 56.3 49.3 34.9 33.6 62.7
    8 58.0 50.0 39.4 33.6 63.9
    16 59.4 50.8 43.0 33.5 64.5
    Flamingo (80B) 0∗0^* 56.3 50.6 31.6 35.0 46.4
    4 63.1 57.4 39.6 36.5 68.6
    8 65.6 57.5 44.8 37.3 70.0
    16 66.8 57.8 48.4 37.6 70.0
    IDEFICS (80B) 0∗0^* 60.0 45.2 36.0 30.9 60.6
    4 63.6 52.4 40.4 34.4 57.8
    8 64.8 55.1 46.1 35.7 58.2
    16 65.4 56.8 48.3 36.3 57.8
    Emu (14B) 0∗0^* 52.9 42.8 34.4 - -
    4 58.4 - 41.3 - -
    8 59.0 - 43.9 - -
    Emu2 (37B) 0 33.5 26.7 40.4 26.4 52.2
    4 67.0 53.2 54.6 48.2 62.4
    8 67.8 54.1 54.7 49.3 65.8
    16 68.8 57.1 57.0 50.3 66.0

    0∗0^* denotes text 2-shot and image 0-shot evaluation following Flamingo conventions. Under 4-, 8-, and 16-shot evaluation, Emu2 (37B) outperforms larger models including Flamingo-80B and IDEFICS-80B on VQAv2, VizWiz, and TextVQA.

  7. Knowl 7 — Visual Question Answering and LMM Benchmark Evaluation of Emu2-Chat

    data/table

    Emu2-Chat was evaluated against generalist large multimodal models on academic visual question answering (VQA), video QA (MSVD, MSRVTT), and holistic large multimodal model (LMM) benchmarks (SEED-Bench, MM-Vet, TouchStone, MMMU).

    Model VQAv2 OKVQA GQA VizWiz TextVQA MSVD MSRVTT SEED MM-Vet TS MMMU
    Flamingo-9B 51.8 44.7 - 28.8 - 30.2 13.7 - - - -
    Flamingo-80B 56.3 50.6 - 31.6 - 35.6 17.4 - - - -
    Kosmos-1 51.0 - - 29.2 - - - - - - -
    Kosmos-2 51.1 - - - - - - 50.0 - - 26.6
    BLIP-2-13B - - 41.0 19.6 42.5 20.3 10.3 46.4 22.4 - -
    InstructBLIP-13B - - 49.5 33.4 50.7 41.2 24.8 - 25.6 552.4 -
    IDEFICS-9B 50.9 38.4 - 35.5 25.9 - - - - - -
    IDEFICS-80B 60.0 45.2 - 36.0 30.9 - - - - - -
    Shikra-13B 77.4* 47.2 - - - - - - - - -
    Qwen-VL-13B-Chat 78.2* 56.6* 57.5* 38.9 61.5* - - 58.2 - 645.2 -
    LLaVA-1.5-13B 80.0* - 63.3* 53.6 61.3 - - 61.6 35.4 - 33.6
    CogVLM 83.4* 58.9* - - 68.1* - - - - 662.6 30.1
    Emu-I 62.0 49.2 46.0 38.3 - 37.0 21.2 - 36.3 - -
    Emu2-Chat 84.9* 64.8* 65.1* 54.9 66.6* 49.0 31.4 62.8 48.5 703.8 34.1
    • indicates that samples from the task's training set were included during training. TS represents TouchStone. On MM-Vet, results reflect the average of five scoring runs. Emu2-Chat achieves top performance across general multimodal understanding, zero-shot video question answering (49.0 on MSVD and 31.4 on MSRVTT despite not being trained on video QA data), and comprehensive multi-discipline reasoning on MMMU (34.1).
  8. Knowl 8 — Referring Expression Comprehension Evaluation of Emu2-Chat

    data/table

    Emu2-Chat was evaluated on visual grounding across the RefCOCO, RefCOCO+, and RefCOCOg datasets:

    Model RefCOCO RefCOCO+ RefCOCOg
    val testA testB val testA testB val test
    OFA-L 79.96 83.67 76.39 68.29 76.00 61.75 67.57 67.58
    Shikra-7B 87.01 90.61 80.24 81.60 87.36 72.12 82.27 82.19
    Shikra-13B 87.83 91.11 81.81 82.89 87.79 74.41 82.64 83.16
    Qwen-VL-7B 89.36 92.26 85.34 83.12 88.25 77.21 85.58 85.48
    Emu2-Chat 90.40 93.88 85.97 87.05 91.43 80.47 87.64 88.11

    Emu2-Chat achieves the highest scores among generalist multimodal models across all splits, exhibiting its largest margin of improvement on RefCOCO+, which restricts descriptions to pure appearance attributes without spatial coordinate cues.

  9. Knowl 9 — Quantitative Evaluation of Zero-Shot Image and Subject-Driven Generation

    data/table

    Emu2-Gen performance was evaluated on zero-shot text-to-image synthesis on the MS-COCO validation set (30,000 samples) and zero-shot single-entity subject-driven generation on DreamBench (3,000 generated images across prompts):

    Zero-Shot Text-to-Image Generation on MS-COCO Validation Set:

    Models CLIP-I ↑\uparrow CLIP-T ↑\uparrow
    Unimodal Generation Models
    MUSE - 0.320
    Imagen - 0.270
    DALL-E 2†^{\dagger} - 0.314
    DALL-E 3†^{\dagger} - 0.320
    SDv1.5 0.667 0.302
    SDXL 0.674 0.310
    Multimodal Generation Models
    GILL 0.684 -
    SEED 0.682 -
    Emu 0.656 0.286
    Emu2-Gen 0.686 0.297

    †^{\dagger} indicates CLIP-T evaluated on 4,096 samples. Emu2 visual autoencoding alone on MS-COCO achieves a CLIP-I reconstruction score of 0.907.

    Single-Entity Subject-Driven Generation on DreamBench:

    Methods DINO ↑\uparrow CLIP-I ↑\uparrow CLIP-T ↑\uparrow
    Real Images (Oracle) 0.774 0.885 -
    Fine-Tuning Methods
    Textual Inversion 0.569 0.780 0.255
    DreamBooth 0.668 0.803 0.305
    BLIP-Diffusion 0.670 0.805 0.302
    Test-Time-Tuning Free Methods
    Re-Imagen* 0.600 0.740 0.270
    SuTI 0.741 0.819 0.304
    BLIP-Diffusion* 0.594 0.779 0.300
    Kosmos-G* (single image input) 0.694 0.847 0.287
    Emu2-Gen* (single image input) 0.766 0.850 0.287
    • denotes zero-shot methods. Emu2-Gen achieves higher subject fidelity (0.766 DINO and 0.850 CLIP-I) than test-time fine-tuning methods (DreamBooth, Textual Inversion) and zero-shot baselines (Kosmos-G, SuTI).

Coverage note — Qualitative visual demonstrations of open-ended conversational examples and specific generation vignettes from the figures were omitted in favor of self-contained architectural definitions, training objectives, and quantitative benchmark tables.

References

  1. 1.Laion-aesthetics. https://laion.ai/blog/laion-aesthetics/, . 4, 5, 2, 3
  2. 2.Laion coco: 600m synthetic captions from laion2b-en. https://laion.ai/blog/laion-coco/, . 4, 2
  3. 3.Laion-high-resolution. https://huggingface.co/datasets/laion/laion-high-resolution, . 5, 3
  4. 4.Sharegpt. https://sharegpt.com/. 5, 2
  5. 5.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022. 1, 5, 6, 8, 3
  6. 6.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023. 6, 7, 8
  7. 7.Shuai Bai, Shusheng Yang, Jinze Bai, Peng Wang, Xingxuan Zhang, Junyang Lin, Xinggang Wang, Chang Zhou, and Jingren Zhou. Touchstone: Evaluating vision-language models by language models. arXiv preprint arXiv:2308.16890, 2023. 6
  8. 8.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738, 2021. 3, 1
  9. 9.James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving image generation with better captions. 2023. 5, 7, 8, 3
  10. 10.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023. 5, 3
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 1, 8
  12. 12.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 8
  13. 13.Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023. 8
  14. 14.Chi Chen, Ruoyu Qin, Fuwen Luo, Xiaoyue Mi, Peng Li, Maosong Sun, and Yang Liu. Position-enhanced visual instruction tuning for multimodal large language models. arXiv preprint arXiv:2308.13437, 2023. 8
  15. 15.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023. 6, 7, 8
  16. 16.Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W Cohen. Re-imagen: Retrieval-augmented text-to-image generator. arXiv preprint arXiv:2209.14491, 2022. 8
  17. 17.Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Rui, Xuhui Jia, Ming-Wei Chang, and William W Cohen. Subject-driven text-to-image generation via apprenticeship learning. arXiv preprint arXiv:2304.00186, 2023. 8
  18. 18.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015. 4, 2
  19. 19.Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov, Jialin Wu, Paul Voigtlaender, Basil Mustafa, Sebastian Goodman, Ibrahim Alabdulmohsin, Piotr Padlewski, et al. Pali-3 vision language models: Smaller, faster, stronger. arXiv preprint arXiv:2310.09199, 2023. 1, 8
  20. 20.Luke Chesser and Timothy Carbone. Unsplash. https://github.com/unsplash/datasets. 5, 3
  21. 21.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 1, 8
  22. 22.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023. 6
  23. 23.Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, and Li Yi. Dreamllm: Synergistic multimodal comprehension and creation. arXiv preprint arXiv:2309.11499, 2023. 8
  24. 24.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 8
  25. 25.Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al. Vision-language pre-training: Basics, recent advances, and future trends. Foundations and Trends® in Computer Graphics and Vision, 14(3–4):163–352, 2022. 1
  26. 26.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020. 3, 2
  27. 27.Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan. Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041, 2023. 4, 7, 8
  28. 28.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 4, 6, 2
  29. 29.Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617, 2018. 6
  30. 30.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 4
  31. 31.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045, 2023. 6, 8
  32. 32.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709, 2019. 4, 6, 2
  33. 33.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 35:26565–26577, 2022. 3
  34. 34.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 787–798, 2014. 5, 6, 2
  35. 35.Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems, 33:2611–2624, 2020. 6
  36. 36.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 5, 3
  37. 37.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. Generating images with multimodal language models. arXiv preprint arXiv:2305.17216, 2023. 7, 8
  38. 38.Hugo Touvron, Lucile Saulnier, Leo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M Rush, Douwe Kiela, et al. Obelics: An open web-scale filtered dataset of interleaved image-text documents. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. 6
  39. 39.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 6
  40. 40.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023. 8
  41. 41.Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. Multimodal foundation models: From specialists to general-purpose assistants. arXiv preprint arXiv:2309.10020, 1(2):2, 2023. 1
  42. 42.Dongxu Li, Junnan Li, and Steven CH Hoi. Blipdiffusion: Pre-trained subject representation for controllable text-to-image generation and editing. arXiv preprint arXiv:2305.14720, 2023. 6, 8
  43. 43.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR, 2022. 8
  44. 44.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 5, 2
  45. 45.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, et al. M3it: A large-scale dataset towards multimodal multilingual instruction tuning. arXiv preprint arXiv:2306.04387, 2023. 4, 2
  46. 46.Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, and Jingdong Wang. Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation. arXiv preprint arXiv:2309.00398, 2023. 4
  47. 47.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014. 7, 8
  48. 48.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023. 4, 6, 3
  49. 49.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 5, 8, 2
  50. 50.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 3, 4
  51. 51.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022. 4
  52. 52.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016. 5, 6, 2
  53. 53.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195–3204, 2019. 4, 6, 2
  54. 54.Midjourney. Midjourney. https://www.midjourney.com. 5, 3
  55. 55.Leonardo Nicoletti and Dina Bass. Humans are biased: Generative ai is even worse. Bloomberg Technology+ Equality. Accessed June, 23:2023, 2023. 1
  56. 56.Xichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng, Wenhu Chen, and Furu Wei. Kosmos-g: Generating images in context with multimodal large language models. arXiv preprint arXiv:2310.02992, 2023. 7, 4
  57. 57.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 3, 5, 6, 8, 1, 2
  58. 58.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023. 1, 3, 4, 8, 2
  59. 59.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 7, 8
  60. 60.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 7
  61. 61.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 8
  62. 62.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 4, 8
  63. 63.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023. 7, 8, 4
  64. 64.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022. 8
  65. 65.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022. 3, 1
  66. 66.Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: a dataset for image captioning with reading comprehension. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 742–758. Springer, 2020. 4, 2
  67. 67.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 4
  68. 68.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 4, 6, 2
  69. 69.Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:2305.16355, 2023. 8
  70. 70.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 3
  71. 71.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multimodality. arXiv preprint arXiv:2307.05222, 2023. 3, 4, 6, 7, 8, 1
  72. 72.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Stanford alpaca: An instruction-following llama model, 2023. 5, 2
  73. 73.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 1, 3
  74. 74.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 4
  75. 75.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pages 23318–23340. PMLR, 2022. 7
  76. 76.Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175, 2023. 8
  77. 77.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023. 6, 7
  78. 78.Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6830–6839, 2023. 8
  79. 79.Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang. Seggpt: Segmenting everything in context. arXiv preprint arXiv:2304.03284, 2023. 8
  80. 80.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022. 8
  81. 81.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022. 8
  82. 82.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017. 6
  83. 83.Charig Yang, Weidi Xie, and Andrew Zisserman. It’s about time: Analog clock reading in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2508–2517, 2022. 5, 2
  84. 84.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023. 8
  85. 85.Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. Ferret: Refer and ground anything anywhere at any granularity. arXiv preprint arXiv:2310.07704, 2023. 8
  86. 86.Lili Yu, Bowen Shi, Ramakanth Pasunuru, Benjamin Muller, Olga Golovneva, Tianlu Wang, Arun Babu, Binh Tang, Brian Karrer, Shelly Sheynin, et al. Scaling autoregressive multimodal models: Pretraining and instruction tuning. arXiv preprint arXiv:2309.02591, 2023. 8
  87. 87.Qiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui, Fan Zhang, Xinlong Wang, and Jingjing Liu. Capsfusion: Rethinking image-text data at scale. arXiv preprint arXiv:2310.20550, 2023. 3, 5, 1
  88. 88.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023. 6
  89. 89.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023. 6
  90. 90.Ao Zhang, Hao Fei, Yuan Yao, Wei Ji, Li Li, Zhiyuan Liu, and Tat-Seng Chua. Transfer visual prompt generator across llms. arXiv preprint arXiv:2305.01278, 2023. 8
  91. 91.Pan Zhang, Xiaoyi Dong Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Hang Yan, et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023. 8
  92. 92.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023. 8
  93. 93.Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun. Llavar: Enhanced visual instruction tuning for text-rich image understanding. arXiv preprint arXiv:2306.17107, 2023. 5, 2
  94. 94.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 8
  95. 95.Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi. Multimodal c4: An open, billion-scale corpus of images interleaved with text. arXiv preprint arXiv:2304.06939, 2023. 3, 1

Citation

MLA
Sun, Q., et al. “Generative Multimodal Models Are In-Context Learners”. arXiv, 2023, http://arxiv.org/abs/2312.13286v2.
APA
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Luo, Z., Wang, Y., Rao, Y., Liu, J., Huang, T., & Wang, X. (2023). Generative Multimodal Models are In-Context Learners. arXiv. http://arxiv.org/abs/2312.13286v2
Chicago
Sun, Q., Y. Cui, X. Zhang, et al. 2023. “Generative Multimodal Models Are In-Context Learners”. arXiv. http://arxiv.org/abs/2312.13286v2.
Harvard
Sun, Q. et al. (2023) “Generative Multimodal Models are In-Context Learners”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.13286v2.
Vancouver
1. Sun Q, Cui Y, Zhang X, et al (2023) Generative Multimodal Models are In-Context Learners. arXiv

BibTeX

@article{sun2023generative,
  title = {Generative Multimodal Models are In-Context Learners},
  author = {Sun, Quan and Cui, Yufeng and Zhang, Xiaosong and Zhang, Fan and Yu, Qiying and Luo, Zhengxiong and Wang, Yueze and Rao, Yongming and Liu, Jingjing and Huang, Tiejun and Wang, Xinlong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.13286v2},
  eprint = {2312.13286}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE