Generative Multimodal Models are In-Context Learners
Quan SunYufeng CuiXiaosong ZhangFan ZhangQiying YuYueze WangYongming RaoJingjing LiuTiejun HuangXinlong Wang
Introduces Emu2, a 37-billion-parameter generative multimodal model trained with a unified autoregressive objective that achieves state-of-the-art in-context learning and controllable generation across diverse visual-language tasks.
Multimodal artificial intelligence systems often require customized architectures and dedicated supervised datasets for every new application, making deployment across diverse real-world tasks difficult and costly. By contrast, humans can learn new concepts or solve complex tasks on the fly using only a few demonstrations or basic instructions. While large language models have mastered this capability in text, multimodal models have lagged behind in flexibly handling mixed sequences of images, video, and text.
The article demonstrates that scaling up a generative multimodal model using a unified learning objective enables strong in-context learning and generalization across both visual understanding and visual content generation. Specifically, it presents and evaluates Emu2, a 37-billion-parameter model designed to operate as a general-purpose multimodal foundation system.
The authors constructed Emu2 by integrating a visual encoder, a multimodal language model backbone, and a visual diffusion decoder, training the entire system with an autoregressive "predict-the-next-element" objective across large-scale text, paired image-video data, and interleaved sequences. The framework was evaluated across two main regimes: few-shot in-context learning with increasing numbers of visual-text examples, and task-specific instruction tuning for dialogue (Emu2-Chat) and controllable image generation (Emu2-Gen). Evaluations encompassed standard visual question answering, referring expression comprehension, complex multimodal reasoning benchmarks, and subject-driven image generation.
Key findings show that Emu2 sets new performance benchmarks while requiring fewer parameters than leading competitors. In few-shot evaluations, Emu2's accuracy consistently improved as context examples grew, outperforming larger 80-billion-parameter models like Flamingo-80B and IDEFICS-80B on benchmarks such as VQAv2 (68.8% vs. 66.8% at 16 shots) and TextVQA (50.3% vs. 37.6%). When instruction-tuned, Emu2-Chat attained state-of-the-art results on comprehensive evaluation suites, scoring 48.5 on MM-Vet and 703.8 on TouchStone, while also leading generalist models in referring expression comprehension on RefCOCO datasets. Additionally, in controllable visual generation, Emu2-Gen achieved industry-leading image and text alignment scores on MS-COCO zero-shot generation and produced superior subject fidelity on DreamBench compared to dedicated tuning-free alternatives.
These results demonstrate that a single, unified generative model can effectively bridge perception and generation, eliminating the need to maintain separate, narrow systems for visual question answering, grounding, and content creation. For organizations, this approach offers substantial efficiency gains and operational simplicity by providing a single interface for diverse multimodal workflows. However, zero-shot performance without prior context examples remains a relative weakness for the base model, emphasizing that contextual prompting or instruction tuning is necessary to unlock optimal performance.
Decision-makers and research teams should consider adopting unified generative architectures as base models for integrated visual and linguistic workflows, while establishing clear instruction-tuning pipelines for specific production tasks. Further work should focus on addressing societal misuse risks, mitigating potential biases inherent in web-scale pretraining datasets, and optimizing inference efficiency before rolling out the system to latency-sensitive or safety-critical operational environments.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Unified-IO introduces a unified sequence-to-sequence framework that processes and generates diverse vision and language formats via a single transformer, establishing the foundation for unified multimodal generative systems like Emu2.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). BLIP-2 establishes the foundational architectural paradigm of interfacing frozen visual encoders with large language model backbones, a technique directly adapted and scaled in Emu2.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). BLIP provides the early blueprint for unifying visual understanding and text generation under shared multimodal objectives.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). MM-Vet introduces the integrated multimodal evaluation benchmark specifically used to assess the conversational and complex reasoning capabilities of Emu2-Chat.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Qwen-VL demonstrates how visual encoding and language generation can be fused for fine-grained multimodal grounding and comprehension, providing a key predecessor for large-scale multimodal chat models.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey systematically formalizes the standard three-component architecture and pretraining-plus-instruction-tuning pipelines utilized to build foundation models such as Emu2.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work demonstrates autoregressive next-token prediction across interleaved text and image tokens, serving as a direct precursor to autoregressive generative multimodal objectives.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides a comprehensive, standardized perceptual and cognitive benchmark essential for evaluating large multimodal foundation models.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o advances unified multimodal understanding and visual generation by combining autoregressive text modelling with discrete diffusion within a single transformer architecture.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). InternVL3 extends large-scale multimodal learning by exploring native joint pre-training paradigms and advanced post-training recipes at up to 78 billion parameters.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL expands on unified multimodal foundation architectures by integrating multi-dimensional positional embeddings and cross-layer visual injection for long-context spatial-temporal reasoning.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). InternVL 2.5 systematically investigates scaling laws across visual encoders, data quality, and test-time reasoning for open-source multimodal models.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision extends unified visual representations across single-image, multi-image, and video reasoning tasks within a single instruction-tuned framework.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). Tuna-2 builds upon unified multimodal architectures by eliminating pretrained vision encoders entirely and operating directly on raw pixel embeddings for both understanding and generation.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni generalizes unified foundation architectures beyond vision and text into real-time omnimodal perception and speech generation.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). UniReal expands controllable visual generation and instruction-guided editing by learning physical dynamics from scalable video data.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). VIEScore demonstrates how multimodal large language models can be applied as explainable, automated evaluators for conditional image synthesis and editing.
