NExT-GPT: Any-to-Any Multimodal LLM

Shengqiong WuHao FeiLeigang QuWei JiTat-Seng Chua

article2024ICML890 citations

Presents an end-to-end multimodal language model that accepts and generates arbitrary combinations of text, image, video, and audio by tuning just one percent of projection parameters between an existing core model and specialized diffusion decoders.

Listen

Current advances in artificial intelligence have produced language models capable of understanding multiple forms of input, such as images, audio, and video. However, most existing systems remain constrained because they only interpret multimodal inputs and output purely textual responses, or they rely on chained pipelines that pass discrete text prompts to external tools. These pipeline systems frequently introduce noise, accumulate transmission errors, and fail to capture complex spatial or numerical instructions. To overcome these limitations, the article introduces NExT-GPT, an end-to-end multimodal artificial intelligence system capable of receiving and generating arbitrary combinations of text, image, video, and audio content.

To build this system efficiently without prohibitive training costs, the authors connected established, frozen foundation models using lightweight trainable adapters. They utilized a unified feature encoder to process multimodal inputs, a central language model for core reasoning and instruction formulation, and off-the-shelf diffusion models to synthesize media outputs. Rather than passing text strings between modules, the central model generates specialized modality signal tokens that directly instruct the downstream diffusion generators. The system was trained using a lightweight alignment strategy paired with a newly curated instruction dataset comprising 5,000 multi-turn, multi-modal dialogue samples designed to teach the model how to switch seamlessly between modalities.

Empirical evaluations show that NExT-GPT achieves competitive or superior performance across multimodal perception, question answering, and media generation tasks. By training only the projection adapters and fine-tuning a small fraction of the central language model parameters—amounting to roughly 1% of total system parameters—the architecture dramatically reduces computational overhead. In comparative tests against pipeline-based baselines, the end-to-end framework scored higher in instruction-following fidelity, logical coherence, and output quality, particularly when resolving complex spatial relationships and object counts that pipeline methods failed to represent accurately.

These findings demonstrate that end-to-end multimodal reasoning and generation can be achieved efficiently without training large models from scratch. For decision-makers and developers, this modular adapter-based design significantly reduces computational resource demands, accelerates deployment timelines, and establishes a scalable foundation for universal artificial intelligence interfaces. However, stakeholders should note that the system's output quality is fundamentally bounded by the capabilities of the underlying base models, and small fine-tuning data volumes can lead to occasional hallucinations or suboptimal media generation, especially for complex video sequences.

Going forward, the authors recommend expanding the framework to support additional modalities such as 3D visual data and document tables, evaluating diverse and larger language model backbones, and scaling up the instruction tuning dataset. Organizations evaluating this technology should pilot it in non-critical interactive domains before deploying it to high-stakes environments, ensuring adequate human oversight and alignment safeguards are maintained.

Cover for NExT-GPT: Any-to-Any Multimodal LLM

Abstract

While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, image, video, and audio. By leveraging the existing well-trained high-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training but also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibilities of building a unified AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Project website: https://next-gpt.github.io/

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Overall Architecture
  • 4. Lightweight Multimodal Alignment Learning
  • 4.1. Encoding-side LLM-centric Multimodal Alignment
  • 4.2. Decoding-side Instruction-following Alignment
  • 5. Modality-switching Instruction Tuning
  • 5.1. Instruction Tuning
  • 5.2. Instruction Dataset
  • 6. Experiments
  • 6.1. Main Results
  • 6.2. In-depth Analysis
  • 7. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Potential Limitation and Future work
  • B. Full Related Work
  • C. Implementation Details
  • C.1. Detailed Input Projection Layer
  • C.2. Model Training
  • C.3. Detailed Dataset
  • C.4. Multimodal IT Datasets Comparison
  • C.5. Training Recipes
  • C.6. Inference Process
  • D. Additional Experiments
  • D.1. Additional Multimodal Comprehension and Generation Results
  • D.2. Human Evaluation on Complex Any-to-any QA
  • D.3. Case Study on Pipeline-style vs. End-to-end Unification
  • D.4. Example Demonstrations

Knowls

  1. Knowl 1 — NExT-GPT Any-to-Any MM-LLM Architecture

    model/method

    NExT-GPT is an end-to-end, general-purpose multimodal large language model (MM-LLM) that accepts and generates arbitrary combinations of text, image, video, and audio. The system is organized into three tiers:

    1. Multimodal Encoding Stage: A unified multimodal encoder, ImageBind (1.2B parameters), encodes raw inputs across image, audio, and video modalities into patch-level representations. An input projection layer (28M parameters) maps these patch representations into semantic concept tokens compatible with the LLM embedding space.

    2. LLM Understanding and Reasoning Stage: A core autoregressive language model, Vicuna-7B (v0), performs cross-modal semantic understanding and multi-step reasoning. Along with standard discrete textual tokens, the LLM outputs dedicated modality signal tokens that instruct downstream diffusion decoders when and what content to synthesize:

      • Image signal tokens: [IMGi]\text{[IMG}_i\text{]} for i∈{0,…,4}i \in \{0, \dots, 4\} (5 tokens total)
      • Audio signal tokens: [AUDi]\text{[AUD}_i\text{]} for i∈{0,…,8}i \in \{0, \dots, 8\} (9 tokens total)
      • Video signal tokens: [VIDi]\text{[VID}_i\text{]} for i∈{0,…,24}i \in \{0, \dots, 24\} (25 tokens total)
    3. Multimodal Generation Stage: Transformer-based output projection layers (31M parameters for image, 31M for audio, 32M for video) project the hidden states of the modality signal tokens into conditional representations. These representations are routed to off-the-shelf latent diffusion decoders:

      • Image synthesis: Stable Diffusion (SD-v1.5, 1.3B parameters)
      • Video synthesis: Zeroscope (v2-576w, 1.8B parameters)
      • Audio synthesis: AudioLDM (l-full, 975M parameters)

    During training, the pre-trained ImageBind encoder and diffusion decoders remain completely frozen, while Vicuna is adapted via LoRA (33M parameters), requiring optimization of only ≈155M\approx 155\text{M} parameters (≈1%\approx 1\% of the total 12.43B12.43\text{B} parameters).

  2. Knowl 2 — Hierarchical Grouping Input Projection Mechanism

    model/method

    To bridge patch-level grid features from multimodal encoders (such as ImageBind) and textual semantic concept tokens in the LLM, NExT-GPT employs a multi-stage hierarchical grouping mechanism rather than direct linear projection.

    Given input patch tokens X1={xi}i=1N1X^1 = \{x_i\}_{i=1}^{N_1} from modality ∗∈{image,audio,video}* \in \{\text{image}, \text{audio}, \text{video}\}, the grouping module performs LL sequential grouping stages. At each stage l∈{1,…,L}l \in \{1, \dots, L\}:

    1. A set of MlM_l learnable concept tokens Cl={cj}j=1MlC^l = \{c_j\}_{j=1}^{M_l} is concatenated with the stage input XlX^l and passed through Transformer layers: [C^l,X^l]=Transformer([Cl;Xl])[\hat{C}^l, \hat{X}^l] = \text{Transformer}([C^l; X^l]) where [⋅;⋅][\cdot; \cdot] denotes sequence concatenation.

    2. A soft assignment similarity matrix AlA^l between updated concept queries C^l\hat{C}^l and patch representations X^l\hat{X}^l is computed via Gumbel-Softmax: Al=Softmax(Norm(C^l)⋅Norm(X^l)⊤+Gτ)A^l = \text{Softmax}\left(\frac{\text{Norm}(\hat{C}^l) \cdot \text{Norm}(\hat{X}^l)^\top + G}{\tau}\right) where GG represents i.i.d. noise sampled from a standard Gumbel(0, 1) distribution, Norm(⋅)\text{Norm}(\cdot) denotes feature normalization, and τ\tau is a learnable temperature parameter.

    3. A hard assignment matrix A^l\hat{A}^l is derived using a straight-through gradient estimator to preserve differentiability during backpropagation: A^l=Onehot(argmax⁡(Al))+Al−Sg⁡(Al)\hat{A}^l = \text{Onehot}(\operatorname{argmax}(A^l)) + A^l - \operatorname{Sg}(A^l) where Sg⁡(⋅)\operatorname{Sg}(\cdot) denotes the stop-gradient operator.

    4. The features are aggregated and projected to form the inputs for the subsequent stage: Xl+1=C^l+MLP(A^l,X^l)X^{l+1} = \hat{C}^l + \text{MLP}(\hat{A}^l, \hat{X}^l)

    After LL stages, the final concept tokens XL∈RML×dX^L \in \mathbb{R}^{M_L \times d} are fed into the LLM as the visual/audio/video prefix prompt.

  3. Knowl 3 — Three-Stage Training Pipeline of NExT-GPT

    model/method

    NExT-GPT is trained through a three-stage progressive alignment and instruction-tuning protocol:

    1. Stage 1: Encoding-side LLM-centric Multimodal Alignment. The objective is to align multimodal encoder features into the LLM token embedding space. The input grouping projection layer is trained on an XX-to-text captioning task using paired data (X∈{image,video,audio}X \in \{\text{image}, \text{video}, \text{audio}\}). The LLM and ImageBind encoder remain frozen. Training minimizes the standard autoregressive cross-entropy loss: Lalign-enc=−∑t=1Tlog⁡P(wt∣w<t,XL)\mathcal{L}_{\text{align-enc}} = - \sum_{t=1}^T \log P(w_t \mid w_{<t}, X^L)

    2. Stage 2: Decoding-side Instruction-Following Alignment. The output projection layers (4-layer Transformer encoder-decoder networks with hidden size 512 and 4 attention heads) are aligned to project LLM modal signal tokens into diffusion conditioning spaces. The diffusion backbones and LLM remain frozen. Training is supervised using paired modal captions and minimizes three joint losses: Lalign-dec=LCE(signal tokens)+λ1∥hsignal−zdiff_text∥22+λ2Ldenoise\mathcal{L}_{\text{align-dec}} = \mathcal{L}_{\text{CE}}(\text{signal tokens}) + \lambda_1 \| h_{\text{signal}} - z_{\text{diff\_text}} \|_2^2 + \lambda_2 \mathcal{L}_{\text{denoise}} where Ldenoise\mathcal{L}_{\text{denoise}} is the conditional latent diffusion denoising score-matching loss, and zdiff_textz_{\text{diff\_text}} is the representation produced by the diffusion model's native text encoder.

    3. Stage 3: End-to-End Instruction Tuning (MosIT). Both the input projection layers, output projection layers, and the LLM backbone (fine-tuned via LoRA) are optimized end-to-end on multi-turn multimodal dialogue datasets (MosIT, T2M, LLaVA-150K, VideoChat, cleaned-Alpaca) using a joint language modeling cross-entropy loss and diffusion generation alignment loss.

  4. Knowl 4 — NExT-GPT System Configuration and Parameter Distribution

    data/table

    NExT-GPT connects frozen pre-trained perception and synthesis models with lightweight learnable adapters. The table below details the component modules, parameter counts, and trainability status:

    Stage / Subsystem Module Name Total Param. Trainable Param. Status
    Multimodal Encoder ImageBind 1.2B 0 Frozen
    Input Projection Grouping Layers 28M 28M Trainable
    Core LLM Vicuna-7B-v0 (LoRA) 7.0B 33M LoRA Tuned
    Image Output Projection Transformer Layer 31M 31M Trainable
    Audio Output Projection Transformer Layer 31M 31M Trainable
    Video Output Projection Transformer Layer 32M 32M Trainable
    Image Diffusion Stable Diffusion v1.5 1.3B 0 Frozen
    Audio Diffusion AudioLDM (l-full) 975M 0 Frozen
    Video Diffusion Zeroscope (v2-576w) 1.8B 0 Frozen
    Total System 12.43B 155M ( 1%)

    Only 155M parameters (155M/[155M+12.275B]≈1.25%155\text{M} / [155\text{M} + 12.275\text{B}] \approx 1.25\%) require gradient updates across all training stages.

  5. Knowl 5 — Modality-Switching Instruction Tuning (MosIT) Dataset

    model/method

    To address the absence of datasets that support bidirectional, multi-turn multimodal interactions across text, image, video, and audio, NExT-GPT introduces the Modality-Switching Instruction Tuning (MosIT) dataset.

    Key characteristics of MosIT compared to prior multimodal IT datasets:

    • Modality Support: Covers inputs and outputs across all four modalities (T+I+A+V→T+I+A+VT+I+A+V \to T+I+A+V), whereas prior datasets (e.g., LLaVA, VideoChat, InstructBLIP, PandaGPT) restrict responses strictly to text (T+X→TT+X \to T).
    • Dialogue Structure: Comprises 5,000 multi-turn dialogues averaging 4.8 turns (3--7 QA pairs per conversation), featuring dynamic alternating modality switches between human and machine.
    • Construction Process: Seed dialogue templates spanning >100>100 topics/keywords were expanded via GPT-4 to generate multi-turn reasoning and planning scenarios. Multimodal assets (images, audios, videos) were retrieved and paired from external sources (YouTube, Google, Flickr) and generative models (Midjourney, Stable Diffusion XL), followed by human quality verification.

    In addition, a Text-to-Multimodal (T2M) dataset containing 15,000 instances was constructed using GPT-4 wrapped around CC3M, WebVid, and AudioCaps captions.

  6. Knowl 6 — Zero-Shot and Supervised Multimodal Perception Performance

    empirical result

    NExT-GPT was evaluated on image, video, and audio comprehension benchmarks against state-of-the-art multimodal LLMs:

    Model Image Captioning (CIDEr ↑\uparrow) Image QA (Accuracy ↑\uparrow) Comprehensive
    NoCaps Flickr30K COCO VQAv2 VizWiz OKVQA MMB SEED
    InstructBLIP 123.1 82.4 102.2 - 33.4 33.9 36.0 -
    LLaVA 120.7 82.7 - - - - 36.2 -
    mPLUG-Owl 117.0 80.3 119.3 - 39.0 - 46.6 34.0
    Emu - - 117.7 40.0 35.4 34.7 - -
    DREAMLLM - - 115.4 56.6 45.8 44.3 49.9 -
    Video-LLaVA - - - 74.7 48.1 - 60.9 -
    NExT-GPT 123.7 84.5 124.9 66.7 48.4 52.1 58.0 57.5

    On video and audio perception benchmarks, NExT-GPT achieves:

    • Video QA: 64.5 on MSVD-QA, 61.4 on MSRVTT-QA, and 50.7 on NExTQA (surpassing Video-LLaMA at 51.6 / 29.6 and Emu at 32.4 / 14.0 / 6.8).
    • Video Captioning (MSR-VTT): 76.2 CIDEr (fine-tuned), outperforming CoDi (74.4) and UIO-2XXL (48.8).
    • Audio Captioning (AudioCaps): 81.3 CIDEr / 0.534 SPIDEr, outperforming CoDi (78.9 CIDEr / 0.480 SPIDEr) and AudioCaps baseline (0.593 CIDEr / 0.369 SPIDEr).
  7. Knowl 7 — Text-Conditioned Multimodal Generation Performance

    empirical result

    NExT-GPT was evaluated on text-to-image (MS COCO), text-to-video (MSRVTT), and text-to-audio (AudioCaps) generation benchmarks against specialized generative systems and multimodal LLMs:

    Model Image FID ↓\downarrow Audio FAD ↓\downarrow Audio FD ↓\downarrow Video CLIPSIM ↑\uparrow
    Stable Diffusion 1.5 11.21 - - -
    AudioLDM-L - - 23.31 -
    CoDi 11.26 1.80 22.90 28.90
    GILL-8B (zero-shot) 12.20 - - -
    Emu-13B (zero-shot) 11.66 - - -
    UIO-2XXL 13.39 2.64 - -
    NExT-GPT 10.07 1.68 23.25 31.97
    NExT-GPT (zero-shot) 11.18 1.74 - 30.96

    NExT-GPT matches or outperforms specialized single-modality and composable diffusion models (such as CoDi) while uniquely providing unified multi-turn conversation and reasoning capabilities across all four output modalities.

  8. Knowl 8 — Ablation of Input Projection Architecture Designs

    empirical result

    To evaluate the grouping mechanism in aligning ImageBind grid representations with LLM semantics, NExT-GPT was evaluated against alternative input projection architectures across image QA, video QA, and audio captioning benchmarks:

    Projection Architecture Image QA Video QA Audio Captioning
    VQAv2 VizWiz MSVD-QA MSRVTT-QA AudioCaps (CIDEr)
    Grouping Mechanism (NExT-GPT) 66.7 48.4 64.5 61.4 81.3
    w/ Linear Layer 63.8 45.4 60.8 57.1 77.4
    w/ Q-Former + Linear Layer 65.1 46.9 63.4 58.1 79.7

    The hierarchical grouping mechanism achieves superior performance over direct linear projection (e.g., +2.9+2.9 on VQAv2, +3.7+3.7 on MSVD-QA, +3.9+3.9 CIDEr on AudioCaps) and Q-Former projection, indicating that aggregating patch-level features into discrete concept-level tokens reduces semantic mismatch with the LLM.

  9. Knowl 9 — Optimal Modality Signal Token Counts for Diffusion Synthesis

    empirical result

    Varying the number of modality signal tokens emitted by the LLM directly influences generation fidelity across modalities:

    • Text-to-Image Generation (FID ↓\downarrow on COCO): Evaluated at token counts {1,2,4,8}\{1, 2, 4, 8\}. The FID score improves from ≈12.3\approx 12.3 at 1 token, reaching an optimal minimum of 10.0710.07 at 4 tokens, before degrading slightly at 8 tokens (>11.0>11.0).

    • Text-to-Audio Generation (FAD ↓\downarrow on AudioCaps): Evaluated at token counts {2,4,8,16}\{2, 4, 8, 16\}. The FAD score decreases steadily from ≈3.4\approx 3.4 at 2 tokens to a minimum of 1.681.68 at 8 tokens, before increasing at 16 tokens (>2.5>2.5).

    • Text-to-Video Generation (FID ↓\downarrow on MSR-VTT): Evaluated at token counts {4,8,16,24,32}\{4, 8, 16, 24, 32\}. Video synthesis requires greater representational capacity due to temporal dynamics, achieving its optimal FID (12.6912.69) at 24 tokens, compared to 16.116.1 at 4 tokens and 13.813.8 at 32 tokens.

    Thus, NExT-GPT fixes signal token counts to 4 for images, 8 for audio, and 24 for video.

  10. Knowl 10 — End-to-End Signal Tokens vs. Intermediate Caption Pipelines

    empirical result

    In human evaluation on 100 complex instructions requiring implicit visual reasoning (scored on a 1--100 scale), NExT-GPT was compared against pipeline systems (HuggingGPT, Visual-ChatGPT) and a text-mediated variant (NExT-GPT-caption) that passes generated discrete text captions to diffusion models instead of soft signal token representations:

    • Instruction Following: NExT-GPT scored ≈84.5\approx 84.5, outperforming NExT-GPT-caption (≈78.5\approx 78.5), Visual-ChatGPT (≈76.0\approx 76.0), and HuggingGPT (≈71.0\approx 71.0).
    • Rationality: NExT-GPT scored ≈82.0\approx 82.0, outperforming NExT-GPT-caption (≈75.0\approx 75.0), Visual-ChatGPT (≈73.0\approx 73.0), and HuggingGPT (≈68.0\approx 68.0).
    • Visual Quality: NExT-GPT scored ≈83.0\approx 83.0, outperforming NExT-GPT-caption (≈76.0\approx 76.0), Visual-ChatGPT (≈75.0\approx 75.0), and HuggingGPT (≈69.5\approx 69.5).

    Passing continuous signal token embeddings directly to diffusion conditioners circumvents the information bottleneck of discrete textual captions, preserving complex non-linguistic semantics such as precise object numeration and fine-grained spatial-relational layouts.

  11. Knowl 11 — Limitations and Scope of NExT-GPT

    limitation

    The NExT-GPT framework has four primary limitations identified by the authors:

    1. Modality and Task Scope: The current implementation supports four core modalities (text, image, video, audio). It lacks native support for 3D vision, web pages, heat maps, tables, figures, and specialized vision-centric tasks (such as object detection, dense segmentation, visual grounding, and tracking).
    2. LLM Scale and Diversity: The core reasoning agent is restricted to the 7B-parameter Vicuna model; larger foundation models and alternative architectures remain unvalidated.
    3. Generative Model Bottleneck and Hallucinations: Synthesis fidelity is bounded by the capabilities of the frozen pre-trained diffusion decoders (SD-v1.5, Zeroscope, AudioLDM). The system is susceptible to cross-modal hallucinations or low-fidelity generation on intricate instructions without retrieval augmentation.
    4. Instruction Tuning Data Volume: The MosIT dataset contains 5,000 instances, leaving room for expansion in diversity and coverage of multi-turn edge cases.

Coverage note — None. All primary architectural components, mathematical formulations, training stages, empirical benchmark comparisons, dataset contributions, and ablation studies have been included as knowls.

References

  1. 1.Agrawal, H., Anderson, P., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., and Lee, S. nocaps: novel object captioning at scale. In Proceedings of the ICCV, pp. 8947–8956, 2019.
  2. 2.Alayrac, J., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K. Flamingo: a visual language model for few-shot learning. In Proceedings of the NeurIPS, 2022.
  3. 3.An, J., Zhang, S., Yang, H., Gupta, S., Huang, J., Luo, J., and Yin, X. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. CoRR, abs/2304.08477, 2023.
  4. 4.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the CVPR, pp. 6077–6086, 2018.
  5. 5.Avrahami, O., Fried, O., and Lischinski, D. Blended latent diffusion. ACM Trans. Graph., 42(4):149:1–149:11, 2023.
  6. 6.Bain, M., Nagrani, A., Varol, G., and Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the ICCV, pp. 1708–1718, 2021.
  7. 7.Bashiri, M., Walker, E. Y., Lurz, K., Jagadish, A., Muhammad, T., Ding, Z., Ding, Z., Tolias, A. S., and Sinz, F. H. A flow-based latent state generative model of neural population responses to natural images. In Proceedings of the NeurIPS, pp. 15801–15815, 2021.
  8. 8.Brock, A., Donahue, J., and Simonyan, K. Large scale GAN training for high fidelity natural image synthesis. In Proceedings of the ICLR, 2019.
  9. 9.Cerspense. Zeroscope: Diffusion-based text-to-video synthesis. 2023. URL https://huggingface.co/cerspense.
  10. 10.Ceylan, D., Huang, C. P., and Mitra, N. J. Pix2video: Video editing using image diffusion. CoRR, abs/2303.12688, 2023.
  11. 11.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 902023.
  12. 12.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V. Y., Huang, Y., Dai, A. M., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models, 2022.
  13. 13.Couairon, G., Verbeek, J., Schwenk, H., and Cord, M. Diffedit: Diffusion-based semantic image editing with mask guidance. In Proceedings of the ICLR, 2023.
  14. 14.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. C. H. Instructblip: Towards general-purpose vision-language models with instruction tuning. CoRR, abs/2305.06500, 2023.
  15. 15.Dess`ı, R., Bevilacqua, M., Gualdoni, E., Rakotonirina, N. C., Franzon, F., and Baroni, M. Cross-domain image captioning with discriminative finetuning. In Proceedings of the CVPR, pp. 6935–6944, 2023.
  16. 16.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., and Tang, J. Cogview: Mastering text-to-image generation via transformers. In Proceedings of the NeurIPS, pp. 19822–19835, 2021.
  17. 17.Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., and Yi, L. Dreamllm: Synergistic multimodal comprehension and creation. CoRR, abs/2309.11499, 2023.
  18. 18.Dorkenwald, M., Milbich, T., Blattmann, A., Rombach, R., Derpanis, K. G., and Ommer, B. Stochastic image-to-video synthesis using cinns. In Proceedings of the CVPR, pp. 3742–3753, 2021.
  19. 19.Fan, W., Chen, Y., Chen, D., Cheng, Y., Yuan, L., and Wang, Y. F. Frido: Feature pyramid diffusion for complex scene image synthesis. CoRR, abs/2208.13753, 2022.
  20. 20.Feng, W., He, X., Fu, T., Jampani, V., Akula, A. R., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y. Training-free structured diffusion guidance for compositional text-to-image synthesis. CoRR, abs/2212.05032, 2022.
  21. 21.Feng, W., Zhu, W., Fu, T., Jampani, V., Akula, A. R., He, X., Basu, S., Wang, X. E., and Wang, W. Y. Layoutgpt: Compositional visual planning and generation with large language models. CoRR, abs/2305.15393, 2023.
  22. 22.Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. CoRR, abs/2208.01618, 2022.
  23. 23.Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J., and Parikh, D. Long video generation with time-agnostic VQGAN and time-sensitive transformer. In Proceedings of the ECCV, pp. 102–118, 2022.
  24. 24.Ge, Y., Ge, Y., Zeng, Z., Wang, X., and Shan, Y. Planting a SEED of vision in large language model. CoRR, abs/2307.08041, 2023.
  25. 25.Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. Audio set: An ontology and human-labeled dataset for audio events. In Proceedings of the ICASSP, pp. 776–780, 2017.
  26. 26.Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. Imagebind: One embedding space to bind them all. CoRR, abs/2305.05665, 2023.
  27. 27.Gontier, F., Serizel, R., and Cerisara, C. Automated audio captioning by fine-tuning BART with audioset tags. In Proceedings of the DCASE, pp. 170–174, 2021.
  28. 28.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the CVPR, pp. 6325–6334, 2017.
  29. 29.Gu, X., Chen, G., Wang, Y., Zhang, L., Luo, T., and Wen, L. Text with knowledge graph augmented transformer for video captioning. In Proceedings of the CVPR, pp. 18941–18951, 2023.
  30. 30.Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the CVPR, pp. 3608–3617, 2018.
  31. 31.Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D. Prompt-to-prompt image editing with cross-attention control. In Proceedings of the ICLR, 2023.
  32. 32.Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. CoRR, abs/2205.15868, 2022.
  33. 33.Hoogeboom, E., Nielsen, D., Jaini, P., Forre, P., and Welling, M. Argmax flows and multinomial diffusion: Towards non-autoregressive language models. CoRR, 2021.
  34. 34.Hsu, W., Bolte, B., Tsai, Y. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE ACM Trans. Audio Speech Lang. Process., 29:3451–3460, 2021.
  35. 35.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. In Proceedings of the ICLR, 2022.
  36. 36.Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X., and Zhao, Z. Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models. In Proceedings of the ICML, pp. 13916–13932, 2023a.
  37. 37.Huang, R., Li, M., Yang, D., Shi, J., Chang, X., Ye, Z., Wu, Y., Hong, Z., Huang, J., Liu, J., Ren, Y., Zhao, Z., and Watanabe, S. Audiogpt: Understanding and generating speech, music, sound, and talking head. CoRR, abs/2304.12995, 2023b.
  38. 38.Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O. K., Patra, B., Liu, Q., Aggarwal, K., Chi, Z., Bjorck, J., Chaudhary, V., Som, S., Song, X., and Wei, F. Language is not all you need: Aligning perception with language models. CoRR, abs/2302.14045, 2023c.
  39. 39.Huang, W., Tu, S., and Xu, L. Pfb-diff: Progressive feature blending diffusion for text-driven image editing. CoRR, abs/2306.16894, 2023d.
  40. 40.Karpathy, A. and Fei-Fei, L. Deep visual-semantic alignments for generating image descriptions. IEEE Trans. Pattern Anal. Mach. Intell., 39(4):664–676, 2017.
  41. 41.Karras, J., Holynski, A., Wang, T., and Kemelmacher-Shlizerman, I. Dreampose: Fashion image-to-video synthesis via stable diffusion. CoRR, abs/2304.06025, 2023.
  42. 42.Kim, C. D., Kim, B., Lee, H., and Kim, G. Audiocaps: Generating captions for audios in the wild. In Proceedings of the NAACL, pp. 119–132, 2019.
  43. 43.Kim, E., Kim, J., Oh, Y., Kim, K., Park, M., Sim, J., Lee, J., and Lee, K. Improving audio-language learning with mixgen and multi-level test-time augmentation. CoRR, abs/2210.17143, 2022.
  44. 44.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), Proceedings of the ICLR, 2015.
  45. 45.Koh, J. Y., Fried, D., and Salakhutdinov, R. Generating images with multimodal language models. CoRR, abs/2305.17216, 2023.
  46. 46.Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y. Seed-bench: Benchmarking multimodal llms with generative comprehension. CoRR, abs/2307.16125, 2023a.
  47. 47.Li, B., Zhang, Y., Chen, L., Wang, J., Pu, F., Yang, J., Li, C., and Liu, Z. MIMIC-IT: multi-modal in-context instruction tuning. CoRR, abs/2306.05425, 2023b.
  48. 48.Li, J., Li, D., Savarese, S., and Hoi, S. C. H. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the ICML, pp. 19730–19742, 2023c.
  49. 49.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. Videochat: Chat-centric video understanding. CoRR, abs/2305.06355, 2023d.
  50. 50.Li, L., Yin, Y., Li, S., Chen, L., Wang, P., Ren, S., Li, M., Yang, Y., Xu, J., Sun, X., Kong, L., and Liu, Q. M3it: A large-scale dataset towards multi-modal multilingual instruction tuning. CoRR, abs/2306.04387, 2023e.
  51. 51.Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Proceedings of the ECCV, pp. 121–137, 2020.
  52. 52.Li, Y., Wang, X., Xiao, J., Ji, W., and Chua, T. Invariant grounding for video question answering. In Proceedings of the CVPR, pp. 2918–2927, 2022.
  53. 53.Li, Y., Zhang, C., Yu, G., Wang, Z., Fu, B., Lin, G., Shen, C., Chen, L., and Wei, Y. Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue data. CoRR, abs/2308.10253, 2023f.
  54. 54.Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L. Video-llava: Learning united visual representation by alignment before projection. CoRR, abs/2311.10122, 2023.
  55. 55.Lin, K., Li, L., Lin, C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y., and Wang, L. Swinbert: End-to-end transformers with sparse attention for video captioning. In Proceedings of the CVPR, pp. 17928–17937, 2022.
  56. 56.Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft COCO: common objects in context. In Fleet, D. J., Pajdla, T., Schiele, B., and Tuytelaars, T. (eds.), Proceedings of the ECCV, pp. 740–755, 2014.
  57. 57.Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D. P., Wang, W., and Plumbley, M. D. Audioldm: Text-to-audio generation with latent diffusion models. In Proceedings of the ICML, pp. 21450–21474, 2023a.
  58. 58.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. CoRR, abs/2304.08485, 2023b.
  59. 59.Liu, S., Wang, T., Bau, D., Zhu, J., and Torralba, A. Diverse image generation via self-conditioned gans. In Proceedings of the CVPR, pp. 14274–14283, 2020.
  60. 60.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., Chen, K., and Lin, D. Mmbench: Is your multi-modal model an all-around player? CoRR, abs/2307.06281, 2023c.
  61. 61.Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A. Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. CoRR, abs/2312.17172, 2023.
  62. 62.Maaz, M., Rasheed, H. A., Khan, S. H., and Khan, F. S. Video-chatgpt: Towards detailed video understanding via large vision and language models. CoRR, abs/2306.05424, 2023.
  63. 63.Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. OK-VQA: A visual question answering benchmark requiring external knowledge. In Proceedings of the CVPR, pp. 3195–3204, 2019.
  64. 64.Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J., and Ermon, S. Sdedit: Guided image synthesis and editing with stochastic differential equations. In Proceedings of the ICLR, 2022.
  65. 65.Milewski, V. S. J., Moens, M., and Calixto, I. Are scene graphs good enough to improve image captioning? In Proceedings of the AACL, pp. 504–515, 2020.
  66. 66.Mou, C., Wang, X., Xie, L., Zhang, J., Qi, Z., Shan, Y., and Qie, X. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023.
  67. 67.Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In Proceedings of the ICML, pp. 16784–16804, 2022.
  68. 68.OpenAI. Introducing chatgpt. 2022a.
  69. 69.OpenAI. Gpt-4 technical report. 2022b.
  70. 70.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Proceedings of the NeurIPS, 2022.
  71. 71.Perazzi, F., Pont-Tuset, J., McWilliams, B., Gool, L. V., Gross, M. H., and Sorkine-Hornung, A. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the CVPR, pp. 724–732, 2016.
  72. 72.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Muller, J., Penna, J., and Rombach, R. SDXL: improving latent diffusion models for high-resolution image synthesis. CoRR, abs/2307.01952, 2023.
  73. 73.Qu, L., Wu, S., Fei, H., Nie, L., and Chua, T. Layoutllm-t2i: Eliciting layout guidance from LLM for text-to-image generation. In Proceedings of the ACM MM, pp. 643–654, 2023a.
  74. 74.Qu, L., Wu, S., Fei, H., Nie, L., and Chua, T. Layoutllm-t2i: Eliciting layout guidance from LLM for text-to-image generation. CoRR, abs/2308.05095, 2023b.
  75. 75.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Proceedings of the ICML, pp. 8748–8763, 2021.
  76. 76.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In Proceedings of the ICML, pp. 8821–8831, 2021.
  77. 77.Razavi, A., van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with VQ-VAE-2. In Proceedings of the NeurIPS, pp. 14837–14847, 2019.
  78. 78.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the CVPR, pp. 10674–10685, 2022.
  79. 79.Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. CoRR, abs/2208.12242, 2022.
  80. 80.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the ACL, pp. 2556–2565, 2018.
  81. 81.Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y. Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface. CoRR, abs/2303.17580, 2023.
  82. 82.Shibata, H., Hanaoka, S., Cao, Y., Yoshikawa, M., Takenaga, T., Nomura, Y., Hayashi, N., and Abe, O. Local differential privacy image generation using flow-based deep generative models. CoRR, abs/2212.10688, 2022.
  83. 83.Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y. Make-a-video: Text-to-video generation without text-video data. CoRR, abs/2209.14792, 2022.
  84. 84.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. In Proceedings of the NeurIPS, 2020.
  85. 85.Su, Y., Lan, T., Liu, Y., Liu, F., Yogatama, D., Wang, Y., Kong, L., and Collier, N. Language models can see: Plugging visual controls in text generation. CoRR, abs/2205.02655, 2022.
  86. 86.Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D. Pandagpt: One model to instruction-follow them all. CoRR, abs/2305.16355, 2023.
  87. 87.Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X. Generative pretraining in multimodality. CoRR, abs/2307.05222, 2023.
  88. 88.Tang, Z., Yang, Z., Zhu, C., Zeng, M., and Bansal, M. Any-to-any generation via composable diffusion. CoRR, abs/2305.11846, 2023.
  89. 89.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. 2023. URL https://github.com/tatsu-lab/stanford_alpaca.
  90. 90.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023.
  91. 91.Vahdat, A. and Kautz, J. NVAE: A deep hierarchical variational autoencoder. In Proceedings of the NeurIPS, 2020.
  92. 92.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Proceedings of the NeurIPS, pp. 5998–6008, 2017.
  93. 93.Veaux, C., Yamagishi, J., MacDonald, K., et al. Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit. CSTR, 6:15, 2017.
  94. 94.Voynov, A., Chu, Q., Cohen-Or, D., and Aberman, K. P+: extended textual conditioning in text-to-image generation. CoRR, abs/2303.09522, 2023.
  95. 95.Wang, J., Yang, Z., Hu, X., Li, L., Lin, K., Gan, Z., Liu, Z., Liu, C., and Wang, L. GIT: A generative image-to-text transformer for vision and language. Trans. Mach. Learn. Res., 2022, 2022a.
  96. 96.Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H. OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In Proceedings of the ICML, volume 162, 2022b.
  97. 97.Wang, T., Yi, J., Fu, R., Tao, J., and Wen, Z. Campnet: Context-aware mask prediction for end-to-end text-based speech editing. IEEE ACM Trans. Audio Speech Lang. Process., 30:2241–2254, 2022c.
  98. 98.Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N. Visual chatgpt: Talking, drawing and editing with visual foundation models. CoRR, abs/2303.04671, 2023.
  99. 99.Wu, J. Z., Ge, Y., Wang, X., Lei, W., Gu, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. CoRR, abs/2212.11565, 2022.
  100. 100.Xiao, J., Shang, X., Yao, A., and Chua, T. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the CVPR, pp. 9777–9786, 2021.
  101. 101.Xiao, J., Yao, A., Liu, Z., Li, Y., Ji, W., and Chua, T. Video as conditional graph hierarchy for multi-granular question answering. In Proceedings of the AAAI, pp. 2804–2812, 2022.
  102. 102.Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the ACM MM, pp. 1645–1653, 2017.
  103. 103.Xu, H., Ye, Q., Yan, M., Shi, Y., Ye, J., Xu, Y., Li, C., Bi, B., Qian, Q., Wang, W., Xu, G., Zhang, J., Huang, S., Huang, F., and Zhou, J. mplug-2: A modularized multimodal foundation model across text, image and video. In Proceedings of the ICML, pp. 38728–38748, 2023.
  104. 104.Xu, J., Mei, T., Yao, T., and Rui, Y. MSR-VTT: A large video description dataset for bridging video and language. In Proceedings of the CVPR, pp. 5288–5296, 2016.
  105. 105.Xu, J., Mello, S. D., Liu, S., Byeon, W., Breuel, T. M., Kautz, J., and Wang, X. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the CVPR, pp. 18113–18123, 2022.
  106. 106.Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the CVPR, pp. 1316–1324, 2018.
  107. 107.Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the ICCV, pp. 1666–1677, 2021.
  108. 108.Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D. Diffsound: Discrete diffusion model for text-to-sound generation. IEEE ACM Trans. Audio Speech Lang. Process., 31:1720–1733, 2023.
  109. 109.Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Dan, Y., Zhao, C., Xu, G., Li, C., Tian, J., Qi, Q., Zhang, J., and Huang, F. mplug-docowl: Modularized multimodal large language model for document understanding. CoRR, abs/2307.02499, 2023a.
  110. 110.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., and Huang, F. mplug-owl: Modularization empowers large language models with multimodality. CoRR, abs/2304.14178, 2023b.
  111. 111.Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., Shao, J., and Ouyang, W. LAMM: language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. CoRR, abs/2306.06687, 2023.
  112. 112.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans. Assoc. Comput. Linguistics, 2:67–78, 2014.
  113. 113.Yu, B., Fu, C., Yu, H., Huang, F., and Li, Y. Unified language representation for question answering over text, tables, and images. CoRR, abs/2306.16762, 2023.
  114. 114.Zeng, Z., Zhang, H., Lu, R., Wang, D., Chen, B., and Wang, Z. Conzic: Controllable zero-shot image captioning by sampling-based polishing. In Proceedings of the CVPR, pp. 23465–23476, 2023.
  115. 115.Zhang, A., Fei, H., Yao, Y., Ji, W., Li, L., Liu, Z., and Chua, T. Transfer visual prompt generator across llms. CoRR, abs/2305.01278, 2023a.
  116. 116.Zhang, B., Gu, S., Zhang, B., Bao, J., Chen, D., Wen, F., Wang, Y., and Guo, B. Styleswin: Transformer-based GAN for high-resolution image generation. In Proceedings of the CVPR, pp. 11294–11304, 2022.
  117. 117.Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. CoRR, abs/2305.11000, 2023b.
  118. 118.Zhang, H., Li, X., and Bing, L. Video-llama: An instruction-tuned audio-visual language model for video understanding. CoRR, abs/2306.02858, 2023c.
  119. 119.Zhang, Y., Zhang, R., Gu, J., Zhou, Y., Lipka, N., Yang, D., and Sun, T. Llavar: Enhanced visual instruction tuning for text-rich image understanding. CoRR, abs/2306.17107, 2023d.
  120. 120.Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., and Zha, Z. Object relational graph with teacher-recommended learning for video captioning. In Proceedings of the CVPR, pp. 13275–13285, 2020.
  121. 121.Zhao, B., Wu, B., and Huang, T. SVIT: scaling up visual instruction tuning. CoRR, abs/2307.04087, 2023a.
  122. 122.Zhao, L., Yu, E., Ge, Z., Yang, J., Wei, H., Zhou, H., Sun, J., Peng, Y., Dong, R., Han, C., and Zhang, X. Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning. CoRR, abs/2307.09474, 2023b.
  123. 123.Zhao, Y., Lin, Z., Zhou, D., Huang, Z., Feng, J., and Kang, B. Bubogpt: Enabling visual grounding in multi-modal llms. CoRR, abs/2307.08581, 2023c.
  124. 124.Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., and Gao, J. Regionclip: Region-based language-image pretraining. In Proceedings of the CVPR, pp. 16772–16782, 2022.
  125. 125.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. CoRR, abs/2304.10592, 2023.
  126. 126.Zhu, M., Pan, P., Chen, W., and Yang, Y. DM-GAN: dynamic memory generative adversarial networks for text-to-image synthesis. In Proceedings of the CVPR, pp. 5802–5810, 2019.

Citation

MLA
Wu, S., et al. “NExT-GPT: Any-to-Any Multimodal LLM”. arXiv, 2023, http://arxiv.org/abs/2309.05519v3.
APA
Wu, S., Fei, H., Qu, L., Ji, W., & Chua, T.-S. (2023). NExT-GPT: Any-to-Any Multimodal LLM. arXiv. http://arxiv.org/abs/2309.05519v3
Chicago
Wu, S., H. Fei, L. Qu, W. Ji, and T.-S. Chua. 2023. “NExT-GPT: Any-to-Any Multimodal LLM”. arXiv. http://arxiv.org/abs/2309.05519v3.
Harvard
Wu, S. et al. (2023) “NExT-GPT: Any-to-Any Multimodal LLM”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.05519v3.
Vancouver
1. Wu S, Fei H, Qu L, Ji W, Chua T-S (2023) NExT-GPT: Any-to-Any Multimodal LLM. arXiv

BibTeX

@article{wu2023next,
  title = {NExT-GPT: Any-to-Any Multimodal LLM},
  author = {Wu, Shengqiong and Fei, Hao and Qu, Leigang and Ji, Wei and Chua, Tat-Seng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.05519v3},
  eprint = {2309.05519}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/