AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Jun ZhanJunqi DaiJiasheng YeYunhua ZhouDong ZhangZhigeng LiuXin ZhangRuibin YuanGe ZhangLinyang Li
Presents AnyGPT, an any-to-any multimodal language model that unifies speech, text, images, and music using discrete sequence modeling without altering standard model architectures or training objectives.
Current artificial intelligence systems struggle to seamlessly process and generate content across diverse communication channels, including text, speech, images, and music. While conventional language models demonstrate advanced reasoning in written language, existing multimodal systems often rely on fragmented combinations of separate encoders and decoders. This structural fragmentation creates representational inconsistencies, complicates training and inference, and restricts interactions to simple text outputs or single non-text modalities.
The article evaluates AnyGPT, a unified any-to-any multimodal language model designed to demonstrate that diverse modalities can be processed and generated within a single, standard language model architecture using discrete sequence modeling.
The researchers adopted a data-centric approach by converting continuous non-text inputs—images, speech, and music—into discrete semantic tokens via specialized tokenizers. These tokens were appended to the standard vocabulary of a pre-trained 7-billion-parameter language model backbone without altering its core architecture or autoregressive training objectives. To address the scarcity of multimodal data, the team assembled a large-scale, text-centric alignment pre-training corpus and synthesized a specialized instruction dataset containing 108,000 multi-turn conversations featuring interleaved multimodal inputs and outputs across all four modalities. High-fidelity rendering was achieved through a two-stage process where the core model generates semantic tokens that are subsequently reconstructed into perceptual outputs using non-autoregressive decoders, diffusion models, and voice cloning tools.
The evaluation yielded several key findings regarding the system's cross-modal capabilities under zero-shot testing conditions. In image understanding, the model achieved a 107.5 captioning score on standard benchmarks, performing comparably to larger specialized vision models. In speech recognition, it attained an 8.5% word error rate without targeted fine-tuning, while its text-to-speech module matched human-like speaker characteristics with a 0.77 similarity score using a three-second reference prompt. In image and music generation, the model scored 0.65 and 0.14 on alignment metrics, demonstrating that unified discrete sequence modeling can support arbitrary combinations of inputs and outputs across text, voice, visual, and musical domains.
These findings indicate that organizations can expand language models into generalist multimodal systems purely through data preprocessing rather than complex structural redesigns. By maintaining a standard transformer framework, teams can leverage existing tooling, reduce engineering complexity, and decrease the software risks associated with maintaining separate architectures for different data types. However, because multimodal training yields slightly higher overall loss than narrow unimodal models, generalist systems may experience slight performance trade-offs relative to dedicated single-task tools.
Decision-makers considering multimodal deployment should evaluate whether a unified system meets specific operational thresholds or if critical workflows still require isolated, specialized models. For future development, the source recommends establishing comprehensive evaluation benchmarks for any-to-any interactions, adopting mixture-of-experts architectures to mitigate cross-modal interference, and extending context lengths beyond current constraints, such as the five-second limit on musical generation.
Confidence in these findings is supported by standardized zero-shot evaluations across multiple public datasets. Nevertheless, readers should account for current operational boundaries, including the restricted duration of generated audio, dependence on discrete tokenizer resolution, and the lack of comprehensive industry benchmarks specifically tailored to interleaved, any-to-any multimodal performance.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). It pioneers the unified sequence-to-sequence tokenization approach that converts diverse input and output modalities into a shared discrete vocabulary for transformer-based modeling.
- Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework, Peng Wang et al. (2022). It provides foundational principles for omni-modal sequence-to-sequence learning by standardizing heterogeneous multimodal tasks into discrete tokens within a unified transformer architecture.
- Paper: Cross-Modal Discrete Representation Learning, Alexander H. Liu et al. (2022). It establishes techniques for cross-modal discrete representation learning using quantized codebooks, which underpins the discrete sequence tokenization used across modalities in AnyGPT.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). It introduces visual instruction tuning on machine-generated conversational data, establishing the multimodal alignment and instruction synthesis methodology that AnyGPT extends to any-to-any modalities.
- Paper: MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning, Zhiyang Xu et al. (2023). It demonstrates how formatting diverse tasks into unified sequence-to-sequence multimodal instructions substantially enhances zero-shot generalist capabilities.
- Paper: Pengi: An Audio Language Model for Audio Tasks, Soham Deshmukh et al. (2023). It models audio and speech tasks as generative sequence problems within a unified language model framework, informing AnyGPT's audio integration.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). It introduces lightweight representation alignment to bridge non-text modalities with frozen large language models, providing the core pretext for non-invasive LLM multimodal adaptation.
- Paper: NExT-GPT: Any-to-Any Multimodal LLM, Shengqiong Wu et al. (2024). It advances any-to-any multimodal conversation by connecting language models with diffusion decoders via specialized modality signal tokens for multi-turn text, image, video, and audio synthesis.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). It scales unified omnimodal architecture to massive parameters and real-time streaming speech and visual interaction via an advanced Thinker-Talker framework.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). It builds upon discrete multimodal tokenization by unifying autoregressive text understanding and discrete diffusion image generation within a single transformer.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). It extends generative next-token multimodal modeling to visual diffusion decoders for scaled in-context multimodal understanding and generation.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It explores advanced native joint pre-training paradigms and post-training preference optimization for modern unified multimodal foundation models.
- Paper: AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond, Zixiang Zhou et al. (2024). It applies unified discrete sequence tokenization within large language models to 3D human motion planning and generation.
- Paper: GPT4Point: A Unified Framework for Point-Language Understanding and Generation, Zhangyang Qi et al. (2024). It extends unified multimodal understanding and generation paradigms from 2D media into the 3D point cloud domain.
