AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun ZhanJunqi DaiJiasheng YeYunhua ZhouDong ZhangZhigeng LiuXin ZhangRuibin YuanGe ZhangLinyang Li

article2024ACL233 citations

Presents AnyGPT, an any-to-any multimodal language model that unifies speech, text, images, and music using discrete sequence modeling without altering standard model architectures or training objectives.

Listen

Current artificial intelligence systems struggle to seamlessly process and generate content across diverse communication channels, including text, speech, images, and music. While conventional language models demonstrate advanced reasoning in written language, existing multimodal systems often rely on fragmented combinations of separate encoders and decoders. This structural fragmentation creates representational inconsistencies, complicates training and inference, and restricts interactions to simple text outputs or single non-text modalities.

The article evaluates AnyGPT, a unified any-to-any multimodal language model designed to demonstrate that diverse modalities can be processed and generated within a single, standard language model architecture using discrete sequence modeling.

The researchers adopted a data-centric approach by converting continuous non-text inputs—images, speech, and music—into discrete semantic tokens via specialized tokenizers. These tokens were appended to the standard vocabulary of a pre-trained 7-billion-parameter language model backbone without altering its core architecture or autoregressive training objectives. To address the scarcity of multimodal data, the team assembled a large-scale, text-centric alignment pre-training corpus and synthesized a specialized instruction dataset containing 108,000 multi-turn conversations featuring interleaved multimodal inputs and outputs across all four modalities. High-fidelity rendering was achieved through a two-stage process where the core model generates semantic tokens that are subsequently reconstructed into perceptual outputs using non-autoregressive decoders, diffusion models, and voice cloning tools.

The evaluation yielded several key findings regarding the system's cross-modal capabilities under zero-shot testing conditions. In image understanding, the model achieved a 107.5 captioning score on standard benchmarks, performing comparably to larger specialized vision models. In speech recognition, it attained an 8.5% word error rate without targeted fine-tuning, while its text-to-speech module matched human-like speaker characteristics with a 0.77 similarity score using a three-second reference prompt. In image and music generation, the model scored 0.65 and 0.14 on alignment metrics, demonstrating that unified discrete sequence modeling can support arbitrary combinations of inputs and outputs across text, voice, visual, and musical domains.

These findings indicate that organizations can expand language models into generalist multimodal systems purely through data preprocessing rather than complex structural redesigns. By maintaining a standard transformer framework, teams can leverage existing tooling, reduce engineering complexity, and decrease the software risks associated with maintaining separate architectures for different data types. However, because multimodal training yields slightly higher overall loss than narrow unimodal models, generalist systems may experience slight performance trade-offs relative to dedicated single-task tools.

Decision-makers considering multimodal deployment should evaluate whether a unified system meets specific operational thresholds or if critical workflows still require isolated, specialized models. For future development, the source recommends establishing comprehensive evaluation benchmarks for any-to-any interactions, adopting mixture-of-experts architectures to mitigate cross-modal interference, and extending context lengths beyond current constraints, such as the five-second limit on musical generation.

Confidence in these findings is supported by standardized zero-shot evaluations across multiple public datasets. Nevertheless, readers should account for current operational boundaries, including the restricted duration of generated audio, dependence on discrete tokenizer resolution, and the lack of comprehensive industry benchmarks specifically tailored to interleaved, any-to-any multimodal performance.

Cover for AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Abstract

We introduce AnyGPT, an any-to-any multi-modal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. AnyGPT can be trained stably without any alterations to the current large language model (LLM) architecture or training paradigms. Instead, it relies exclusively on data-level preprocessing, facilitating the seamless integration of new modalities into LLMs, akin to the incorporation of new languages. We build a multimodal text-centric dataset for multimodal alignment pre-training. Utilizing generative models, we synthesize the first large-scale any-to-any multimodal instruction dataset. It consists of 108k samples of multi-turn conversations that intricately interweave various modalities, thus equipping the model to handle arbitrary combinations of multimodal inputs and outputs. Experimental results demonstrate that AnyGPT is capable of facilitating any-to-any multimodal conversation while achieving performance comparable to specialized models across all modalities, proving that discrete representations can effectively and conveniently unify multiple modalities within a language model. Demos are shown in https://junzhan2000.github.io/AnyGPT.github.io/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multimodal Large Language Models
  • 2.2 Multimodal Discretization
  • 3 AnyGPT
  • 3.1 Tokenization
  • 3.2 Language Model Backbone
  • 3.3 Multimodal Generation
  • 4 Multimodal Data
  • 4.1 Pre-training Data
  • 5 Experiment
  • 5.1 Evaluation
  • 5.1.1 Image
  • 5.1.2 Speech
  • 5.1.3 Music
  • 5.2 Example Demonstrations
  • 6 Conclusion
  • Limitations and Future Work
  • Any-to-Any Multimodal LLM Benchmark
  • Acknowledgements
  • References
  • A Pretraining
  • A.1 Data
  • A.2 Pre-training
  • B Instruction Tuning
  • C Evaluation
  • D Prompts for Constructing Multimodal Interleaved Instruction Data
  • E Examples Demonstration

Knowls

  1. Knowl 1 — Unified Discrete Representation Architecture of AnyGPT

    model/method

    AnyGPT is an any-to-any multimodal large language model designed to perceive, understand, reason over, and generate content across text, speech, images, and music within a single autoregressive framework without altering the underlying Transformer architecture. Continuous non-text inputs are tokenized into discrete semantic tokens using modality-specific tokenizers, arranged into interleaved multimodal sequences, and processed by the core language model using standard next-token prediction loss.

    The base architecture is built upon LLaMA-2 7B, pre-trained on 2 TB of text tokens. To integrate non-text modalities, the model expands the vocabulary and resizes the embedding and prediction layers with randomly initialized parameters for each new modality. The total vocabulary size VV is defined as:

    V=∑i=1nViV = \sum_{i=1}^{n} V_i

    where nn denotes the total number of supported modalities and ViV_i represents the vocabulary size of the ii-th modality. All modalities share a unified representational space inside the language model.

  2. Knowl 2 — Modality-Specific Discrete Tokenization in AnyGPT

    model/method

    AnyGPT employs specialized discrete tokenizers to convert non-text continuous signals into discrete token sequences:

    1. Image Tokenization: Uses the SEED tokenizer. A Vision Transformer (ViT) processes a 224×224224 \times 224 RGB image into 16×1616 \times 16 patches, which a Causal Q-Former converts into 32 causal embeddings. An 8,192-entry vector quantization (VQ) codebook discretizes these embeddings into 32 discrete visual tokens per image.
    2. Speech Tokenization: Uses SpeechTokenizer, an encoder-decoder network employing Residual Vector Quantization (RVQ) with 8 hierarchical quantizers of codebook size 1,024 each at a sampling frame rate of 50 Hz. Only the first quantizer layer (capturing semantic content, vocabulary size 1,024) is modeled by the autoregressive LLM (producing 50 tokens per second). Layers 2 through 8 capture paralinguistic acoustic details.
    3. Music Tokenization: Uses Encodec pre-trained on 32 kHz monophonic audio at 50 Hz. It employs 4 RVQ quantizers with a codebook size of 2,048 each, yielding a combined vocabulary size of 8,192 (4×20484 \times 2048). A 5-second music clip is quantized into a 250×4250 \times 4 code matrix and flattened into a causal 1D sequence of 1,000 tokens (200 tokens per second) predicted frame by frame.
    Modality Image Speech Music
    Vocab Size 8192 1024 8192
    Tokens per Sample 32 / image 50 / s 200 / s
    RVQ ✓ ✓
    Input Size 224 × 224 variable duration 5s
  3. Knowl 3 — Two-Stage Hierarchical Multimodal Generation Framework

    model/method

    To prevent computational complexity from exploding with long token sequences required for high-fidelity multimodal synthesis, AnyGPT uses a two-stage generation process separating semantic modeling from perceptual modeling:

    1. Semantic Information Modeling: The autoregressive language model backbone generates discrete semantic tokens aligned across modalities.
    2. Perceptual Information Modeling: Modality-specific non-autoregressive decoders and generative diffusion models convert semantic tokens back into continuous raw perceptual signals:
      • Image Generation: Discrete SEED tokens are mapped via a multi-layer perceptron (MLP) into generation embeddings aligned with the latent space of unCLIP Stable Diffusion, and a UNet decoder synthesizes the output image.
      • Speech Synthesis and Voice Cloning: A non-autoregressive Masked Language Model (SoundStorm, trained on Multilingual LibriSpeech) predicts SpeechTokenizer's acoustic codes (layers 2--8) conditioned on the LLM's semantic tokens and a 3-second reference audio prompt. The SpeechTokenizer neural decoder then decodes the 8-layer RVQ matrix into raw audio, matching the target speaker's voice timbre and emotion.
      • Music Synthesis: The Encodec neural audio decoder directly reconstructs raw 32 kHz audio from the 4-layer discrete music tokens generated by the LLM.
  4. Knowl 4 — AnyInstruct-108k Interleaved Multimodal Instruction Dataset Synthesis

    model/method

    AnyInstruct-108k is a synthetic multimodal instruction dataset designed to train language models on interleaved, multi-turn dialogues across text, speech, images, and music. The dataset is generated through a two-stage synthesis pipeline:

    1. Text-Based Conversation Generation with Multimodal Placeholders:

      • Topic Expansion: 100 brainstormed meta-topics covering audiovisual themes are expanded into 20,000 fine-grained conversational topics using GPT-4.
      • Scenario Construction: GPT-4 generates interactive scenarios from topics using few-shot demonstrations and explicit requirement sampling across user/assistant modality actions (e.g., "the user shares music", "the user asks for images").
      • Chat Script Generation: GPT-4 generates 2-to-3-turn dialogues where non-text content is represented as detailed descriptive tags (e.g., [image: description] and [music: description]), restricting utterances to 5--15 words.
    2. Text-to-Multimodality Conversion:

      • Image descriptions are rendered into images using DALL-E 3.
      • Music descriptions are rendered into audio tracks using MusicGen.
      • Spoken user instructions and assistant voice outputs are synthesized using Microsoft Azure Text-to-Speech API.

    The final dataset consists of 108k multimodal dialogues comprising ~205k images, ~503k voice recordings, and ~113k music tracks, supplemented by 100k speech dialogues rendered from text-only instruction datasets.

  5. Knowl 5 — Text-Centric Pre-Training Data Composition and Sample Construction

    experimental setup

    AnyGPT constructs a text-centric multimodal alignment corpus to bridge diverse modalities via natural language. Modalities with smaller data quantities are oversampled to balance batch composition during pre-training.

    Modality Dataset Description Sample Rate
    Interleaved Image-Text MMC4-core-ff 7.3M documents from Common Crawl 0.05
    Image-Text LAION-2B, LAION-COCO, 300M filtered image-text pairs; 0.30
    JourneyDB, LAION-Aesthetics 4.4M Midjourney pairs
    Speech-Text Multilingual LibriSpeech 44,000-hour English audiobook subset 0.13
    CommonVoice, GigaSpeech 3,000h volunteer recordings; 10,000h audio 0.27
    Music-Text Youtube-Music-1M 100M music-text pairs with GPT-4 captions 0.25
    MusicGen-Synthesis 20k synthesized pairs from AnyInstruct-108k

    For paired data (non-text sequence SS and text TT), GPT-4 creates bidirectional task instructions II. Training instances are wrapped into bidirectional templates:

    • Text-to-Modality: [Human]: {I}. This is input:{T}<eoh>. [AnyGPT]: {S}<eos>.
    • Modality-to-Text: [Human]: {I}.{S}<eoh>. [AnyGPT]: {T}<eos>.

    Samples from the same dataset are concatenated into sequences of maximum length 4,500 tokens to maximize training efficiency.

  6. Knowl 6 — Training Hyperparameters and Decoding Configurations of AnyGPT

    experimental setup

    AnyGPT is trained in two phases: pre-training on bi-modal alignment data followed by instruction fine-tuning on AnyInstruct-108k. Initial pre-training is performed for 81,000 steps, followed by an additional 4,000 steps focusing on high-quality subsets (JourneyDB, LAION-Aesthetics, LAION-COCO, and AnyInstruct-108k).

    Hyperparameter Pre-training Stage Fine-tuning Stage
    Gradient clipping (Global-norm) 1.0 1.0
    Batch size 480 64
    Max length 4500 4500
    Training steps 81000 5000
    Learning rate scheduler cosine cosine
    Peak learning rate 6e-5 2e-5
    Warmup ratio 0.03 0.03
    Optimizer Adam Adam
    Hardware A100 GPUs A100 GPUs

    During inference evaluation, modality-specific decoding strategies are applied:

    Generated Modality Text Image Speech Music
    Decoding Strategy Beam Search Sampling Sampling Sampling
    Beam size 5 – – –
    Top-pp – 0.7 0.7 1.0
    Repetition Penalty 1.0 1.0 1.0 1.15
  7. Knowl 7 — Zero-Shot Cross-Modal Understanding and Generation Benchmark Results

    empirical result

    AnyGPT is evaluated zero-shot across cross-modal perception and generation tasks without task-specific fine-tuning:

    1. Image Understanding (MS-COCO Karpathy split captioning):

      • AnyGPT (8B) achieves a CIDEr score of 107.5, compared to Flamingo 9B (79.4), Flamingo 80B (84.3), Emu 14B (112.4), InstructBLIP 14B (102.2), and SEED-LLaMA 8B (123.6).
    2. Image Generation (MS-COCO 30k validation text-to-image):

      • AnyGPT achieves a CLIPscore of 0.65 based on CLIP-ViT-L, compared to GILL (0.67), Emu (0.66), and SEED-LLaMA (0.69).
    3. Speech Recognition (LibriSpeech test-clean ASR):

      • AnyGPT achieves a Word Error Rate (WER) of 8.5%, compared to Wav2vec 2.0 (2.7%), Whisper Large V2 (2.7%), and human-level performance (5.8%).
    4. Text-to-Speech (VCTK Zero-Shot TTS):

      • Evaluated using Whisper medium for WER and WavLM-TDNN cosine similarity for speaker similarity (SIM) against a 3-second vocal prompt:
    Method WER ↓\downarrow SIM ↑\uparrow
    Ground Truth 1.9 0.93
    VALL-E 7.9 0.75
    USLM 6.5 0.84
    AnyGPT 8.5 0.77
    1. Music Understanding and Generation (MusicCaps benchmark):
      • Evaluated using CLAPscore. On music understanding (music captioning), AnyGPT achieves a CLAPscore of 0.11 (against 0.16 for ground-truth captions). On music generation, AnyGPT achieves a CLAPscore of 0.14, compared to Riffusion (0.19) and Mousai (0.23).
  8. Knowl 8 — Limitations of Discrete Sequence Modeling in AnyGPT

    limitation

    The discrete sequence modeling approach in AnyGPT exhibits several limitations:

    1. Multimodal Loss Gap: Multimodal LLMs trained with discrete representations exhibit higher training loss than unimodal models, leading to performance trade-offs across individual modalities.
    2. Audio Sequence Length Constraints: Due to quadratic attention complexity with sequence length, music generation is restricted to 5-second clips (250 frames×4 codebooks=1000 tokens250 \text{ frames} \times 4 \text{ codebooks} = 1000 \text{ tokens}), limiting practical utility for extended music synthesis.
    3. Tokenizer Quality Ceiling: Compression of continuous visual and auditory signals into low-bitrate discrete tokens introduces information loss and acts as an upper performance bound on comprehension and synthesis fidelity.
    4. Absence of Standardized Any-to-Any Benchmarks: There is a lack of comprehensive, multi-turn, any-to-any benchmarks to evaluate interactive cross-modal reasoning and potential risks across arbitrary modality combinations.

Coverage note — Specific prompt templates used during conversation brainstorming (Figures 5, 6, 7) and qualitative conversational dialogues shown in Appendix E (Figures 8--16) were omitted as their substance is fully represented in the instruction dataset generation and evaluation knowls.

References

  1. 1.2022a. Laion-aesthetics. https://laion.ai/blog/laion-aesthetics/.
  2. 2.2022b. Laion coco: 600m synthetic captions from laion2b-en. https://laion.ai/blog/laion-coco/.
  3. 3.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. ArXiv preprint, abs/2303.08774.
  4. 4.Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matthew Sharifi, Neil Zeghidour, and Christian Havnø Frank. 2023. Musiclm: Generating music from text. ArXiv preprint, abs/2301.11325.
  5. 5.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a visual language model for few-shot learning. ArXiv preprint, abs/2204.14198.
  6. 6.Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al. 2016. Deep speech 2: End-to-end speech recognition in english and mandarin. In International conference on machine learning, pages 173–182. PMLR.
  7. 7.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4218–4222, Marseille, France. European Language Resources Association.
  8. 8.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  9. 9.James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. 2023. Improving image generation with better captions. Computer Science. https://cdn.openai.com/papers/dall-e-3.pdf, 2(3):8.
  10. 10.Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023a. Audiolm: a language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  11. 11.Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi. 2023b. Soundstorm: Efficient parallel audio generation. ArXiv preprint, abs/2305.09636.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  13. 13.Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Zhao You, and Zhiyong Yan. 2021. Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio. In Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021, pages 3670–3674. ISCA.
  14. 14.Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023. Simple and controllable music generation. ArXiv preprint, abs/2306.05284.
  15. 15.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. ArXiv preprint, abs/2305.06500.
  16. 16.Alexandre D’efossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022. High fidelity neural audio compression. ArXiv preprint, abs/2210.13438.
  17. 17.Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jian-Yuan Sun, Hongyu Zhou, Hao-Ran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, and Li Yi. 2023. Dreamllm: Synergistic multimodal comprehension and creation. ArXiv preprint, abs/2309.11499.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  19. 19.Seth* Forsgren and Hayk* Martiros. 2022. Riffusion - Stable diffusion for real-time music generation.
  20. 20.Josh Gardner, Simon Durand, Daniel Stoller, and Rachel M. Bittner. 2023. Llark: A multimodal instruction-following language model for music.
  21. 21.Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan. 2023a. Planting a seed of vision in large language model. ArXiv preprint, abs/2307.08041.
  22. 22.Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li, Xintao Wang, and Ying Shan. 2023b. Making llama see and draw with seed tokenizer. ArXiv preprint, abs/2310.01218.
  23. 23.Rongjie Huang, Jia-Bin Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiaoyue Yin, and Zhou Zhao. 2023. Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models. ArXiv preprint, abs/2301.12661.
  24. 24.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. 2023. Generating images with multimodal language models. ArXiv preprint, abs/2305.17216.
  25. 25.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning.
  26. 26.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European Conference on Computer Vision.
  27. 27.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems, 36.
  28. 28.Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi. 2023. Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. ArXiv preprint, abs/2312.17172.
  29. 29.Microsoft. Microsoft azure text-to-speech api. https://azure.microsoft.com/en-us/products/ai-services/ai-speech.
  30. 30.Junting Pan, Keqiang Sun, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, et al. 2023. Journeydb: A benchmark for generative image understanding. ArXiv preprint, abs/2307.00716.
  31. 31.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210. IEEE.
  32. 32.Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. 2020. MLS: A large-scale multilingual dataset for speech research. In Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020, pages 2757–2761. ISCA.
  33. 33.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  34. 34.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. Robust speech recognition via large-scale weak supervision. ArXiv preprint, abs/2212.04356.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  36. 36.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695.
  37. 37.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer.
  38. 38.Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf. 2023. Moûsai: Text-to-music generation with long-context latent diffusion. ArXiv preprint, abs/2301.11757.
  39. 39.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294.
  40. 40.Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2023a. Generative multimodal models are in-context learners. ArXiv preprint, abs/2312.13286.
  41. 41.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2023b. Generative pretraining in multimodality. ArXiv preprint, abs/2307.05222.
  42. 42.Tianxiang Sun, Xiaotian Zhang, Zhengfu He, Peng Li, Qinyuan Cheng, Xiangyang Liu, Hang Yan, Yunfan Shao, Qiong Tang, Shiduo Zhang, Xingjian Zhao, Ke Chen, Yining Zheng, Zhejian Zhou, Ruixiao Li, Jun Zhan, Yunhua Zhou, Linyang Li, Xiaogui Yang, Lingling Wu, Zhangyue Yin, Xuanjing Huang, Yu-Gang Jiang, and Xipeng Qiu. 2024. Moss: An open conversational large language model. Machine Intelligence Research.
  43. 43.Zineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu, Chenguang Zhu, and Mohit Bansal. 2023a. Codi-2: In-context, interleaved, and interactive any-to-any generation. ArXiv preprint, abs/2311.18775.
  44. 44.Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. 2023b. Any-to-any generation via composable diffusion. ArXiv preprint, abs/2305.11846.
  45. 45.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, abs/2307.09288.
  46. 46.Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural discrete representation learning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6306–6315.
  47. 47.Chengyi Wang, Sanyuan Chen, Yu Wu, Zi-Hua Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. 2023. Neural codec language models are zero-shot text to speech synthesizers. ArXiv preprint, abs/2301.02111.
  48. 48.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. ArXiv preprint, abs/2212.10560.
  49. 49.Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2023. Next-gpt: Any-to-any multimodal llm. ArXiv preprint, abs/2309.05519.
  50. 50.Yusong Wu, K. Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.
  51. 51.Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507.
  52. 52.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023a. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. In Conference on Empirical Methods in Natural Language Processing.
  53. 53.Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu. 2023b. Speechtokenizer: Unified speech tokenizer for speech large language models. ArXiv preprint, abs/2308.16692.
  54. 54.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.
  55. 55.Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi. 2023b. Multimodal c4: An open, billion-scale corpus of images interleaved with text. ArXiv preprint, abs/2304.06939.

Citation

MLA
Zhan, J., et al. “AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling”. arXiv, 2024, http://arxiv.org/abs/2402.12226v5.
APA
Zhan, J., Dai, J., Ye, J., Zhou, Y., Zhang, D., Liu, Z., Zhang, X., Yuan, R., Zhang, G., Li, L., Yan, H., Fu, J., Gui, T., Sun, T., Jiang, Y.-G., & Qiu, X. (2024). AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling. arXiv. http://arxiv.org/abs/2402.12226v5
Chicago
Zhan, J., J. Dai, J. Ye, et al. 2024. “AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling”. arXiv. http://arxiv.org/abs/2402.12226v5.
Harvard
Zhan, J. et al. (2024) “AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.12226v5.
Vancouver
1. Zhan J, Dai J, Ye J, et al (2024) AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling. arXiv

BibTeX

@article{zhan2024anygpt,
  title = {AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling},
  author = {Zhan, Jun and Dai, Junqi and Ye, Jiasheng and Zhou, Yunhua and Zhang, Dong and Liu, Zhigeng and Zhang, Xin and Yuan, Ruibin and Zhang, Ge and Li, Linyang and Yan, Hang and Fu, Jie and Gui, Tao and Sun, Tianxiang and Jiang, Yu-Gang and Qiu, Xipeng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.12226v5},
  eprint = {2402.12226}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/