NExT-GPT: Any-to-Any Multimodal LLM
Shengqiong WuHao FeiLeigang QuWei JiTat-Seng Chua
Presents an end-to-end multimodal language model that accepts and generates arbitrary combinations of text, image, video, and audio by tuning just one percent of projection parameters between an existing core model and specialized diffusion decoders.
Current advances in artificial intelligence have produced language models capable of understanding multiple forms of input, such as images, audio, and video. However, most existing systems remain constrained because they only interpret multimodal inputs and output purely textual responses, or they rely on chained pipelines that pass discrete text prompts to external tools. These pipeline systems frequently introduce noise, accumulate transmission errors, and fail to capture complex spatial or numerical instructions. To overcome these limitations, the article introduces NExT-GPT, an end-to-end multimodal artificial intelligence system capable of receiving and generating arbitrary combinations of text, image, video, and audio content.
To build this system efficiently without prohibitive training costs, the authors connected established, frozen foundation models using lightweight trainable adapters. They utilized a unified feature encoder to process multimodal inputs, a central language model for core reasoning and instruction formulation, and off-the-shelf diffusion models to synthesize media outputs. Rather than passing text strings between modules, the central model generates specialized modality signal tokens that directly instruct the downstream diffusion generators. The system was trained using a lightweight alignment strategy paired with a newly curated instruction dataset comprising 5,000 multi-turn, multi-modal dialogue samples designed to teach the model how to switch seamlessly between modalities.
Empirical evaluations show that NExT-GPT achieves competitive or superior performance across multimodal perception, question answering, and media generation tasks. By training only the projection adapters and fine-tuning a small fraction of the central language model parameters—amounting to roughly 1% of total system parameters—the architecture dramatically reduces computational overhead. In comparative tests against pipeline-based baselines, the end-to-end framework scored higher in instruction-following fidelity, logical coherence, and output quality, particularly when resolving complex spatial relationships and object counts that pipeline methods failed to represent accurately.
These findings demonstrate that end-to-end multimodal reasoning and generation can be achieved efficiently without training large models from scratch. For decision-makers and developers, this modular adapter-based design significantly reduces computational resource demands, accelerates deployment timelines, and establishes a scalable foundation for universal artificial intelligence interfaces. However, stakeholders should note that the system's output quality is fundamentally bounded by the capabilities of the underlying base models, and small fine-tuning data volumes can lead to occasional hallucinations or suboptimal media generation, especially for complex video sequences.
Going forward, the authors recommend expanding the framework to support additional modalities such as 3D visual data and document tables, evaluating diverse and larger language model backbones, and scaling up the instruction tuning dataset. Organizations evaluating this technology should pilot it in non-critical interactive domains before deploying it to high-stakes environments, ensuring adequate human oversight and alignment safeguards are maintained.
- Paper: ImageBind One Embedding Space to Bind Them All, Rohit Girdhar et al. (2023). ImageBind establishes the shared multi-sensory embedding space bridging text, images, video, and audio that NExT-GPT utilizes for its unified multimodal input encoding.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). MiniGPT-4 introduced the foundational strategy of aligning frozen multimodal components to a central LLM backbone using lightweight projection layers.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). Video-LLaMA outlines the adapter-based approach for bridging auditory and temporal visual signals into language models, establishing the input-side multi-modal alignment adapted by NExT-GPT.
- Paper: One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale, Fan Bao et al. (2023). UniDiffuser demonstrates multi-modal conditional generation using diffusion models, providing key conceptual foundations for multi-modal decoding from latent representations.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Video-LLaVA provides important background on unified visual representation and instruction tuning across images and video prior to projection into LLMs.
- Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework, Peng Wang et al. (2022). OFA outlines foundational sequence-to-sequence unified task architectures that motivate all-in-one generalist multimodal systems.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey establishes the standard architectural paradigm—connecting modality encoders to LLMs via lightweight adapters—which NExT-GPT extends to multimodal output synthesis.
- Paper: HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face, Yongliang Shen et al. (2023). HuggingGPT demonstrates the concept of using an LLM controller to coordinate multimodal perception and generation models before end-to-end any-to-any systems were developed.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o progresses beyond adapter-linked diffusion decoders by implementing a single transformer architecture that natively unifies autoregressive understanding and discrete diffusion generation.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni extends end-to-end any-to-any multimodal LLM generation to real-time omnimodal reasoning and streaming speech interaction at massive scale.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). Emu2 scales up unified multimodal generation and in-context learning across interleaved vision-language sequences using integrated visual diffusion decoders.
- Paper: Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners, Yazhou Xing et al. (2024). Seeing and Hearing explores synchronized visual-audio generation using diffusion latent aligners within shared embedding spaces like ImageBind.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). Video-LaVIT advances video understanding and generation within a unified language model framework through decoupled visual-motional tokenization.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). Tuna-2 explores an alternative architecture for unified multimodal understanding and generation by operating directly on raw pixel embeddings rather than pretrained encoders and decoders.
- Paper: VideoPoet: A Large Language Model for Zero-Shot Video Generation, Dan Kondratyuk et al. (2024). VideoPoet generalizes multimodal generation by using an autoregressive language model over discrete visual and audio tokens rather than external diffusion pipelines.
- Paper: AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond, Zixiang Zhou et al. (2024). AvatarGPT applies the principles of unified multimodal understanding and generation to 3D human motion planning and synthesis within foundation models.
