Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal dialogues

Multimodal dialogues are conversational interactions between participants, such as humans and artificial intelligence systems, that incorporate and exchange information across multiple modes of communication rather than relying solely on text. In these multi-turn exchanges, participants can combine, alternate, and interpret various modalities, including spoken language, written text, visual imagery, video, and audio, as both inputs and outputs. Systems supporting multimodal dialogue dynamically track context across conversational turns, align heterogeneous data representations, and generate appropriate multimodal responses, facilitating more versatile and natural communication.

1 item

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu

OrganizationsFudan UniversityMultimodal Art Projection Research CommunityShanghai Artificial Intelligence Laboratory

Why you should read this

Presents AnyGPT, an any-to-any multimodal language model that unifies speech, text, images, and music using discrete sequence modeling without altering standard model architectures or training objectives.

We introduce AnyGPT, an any-to-any multi-modal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. AnyGPT can be trained stably without any alterations to the current large language model (LLM) architecture or training paradigms. Instead, it relies exclusively on data-level preprocessing, facilitating the seamless integration of new modalities into LLMs, akin to the incorporation of new languages. We build a multimodal text-centric dataset for multimodal alignment pre-training. Utilizing generative models, we synthesize the first large-scale any-to-any multimodal instruction dataset. It consists of 108k samples of multi-turn conversations that intricately interweave various modalities, thus equipping the model to handle arbitrary combinations of multimodal inputs and outputs. Experimental results demonstrate that AnyGPT is capable of facilitating any-to-any multimodal conversation while achieving performance comparable to specialized models across all modalities, proving that discrete representations can effectively and conveniently unify multiple modalities within a language model. Demos are shown in https://junzhan2000.github.io/AnyGPT.github.io/.

Added

2026-09-29