Built independently by an author, for readers. Read the story and support ChapterPal

keyword

modality-switching instruction

A modality-switching instruction is a prompt or training directive in multimodal artificial intelligence that requires a model to process inputs in one or more data formats, such as text, images, audio, or video, and transition across modalities to generate outputs in a different or combined format. Typically utilized in multimodal instruction tuning, these instructions train systems to understand cross-modal semantics and dynamically switch between input and output representations during single-turn or multi-turn interactions. By pairing diverse combinations of sensory inputs with different target modalities, modality-switching instructions enable models to perform complex tasks such as synthesizing images or audio from textual descriptions, translating visual data into other media, and alternating communication channels within a coherent conversational flow.

1 item

NExT-GPT: Any-to-Any Multimodal LLM

NExT-GPT: Any-to-Any Multimodal LLM

Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, Tat-Seng Chua

OrganizationsNational University of SingaporeNExT++ Research Center

Why you should read this

Presents an end-to-end multimodal language model that accepts and generates arbitrary combinations of text, image, video, and audio by tuning just one percent of projection parameters between an existing core model and specialized diffusion decoders.

While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, image, video, and audio. By leveraging the existing well-trained high-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training but also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibilities of building a unified AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Project website: https://next-gpt.github.io/

Added

2026-09-28