Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Han ZhaoMin ZhangWei ZhaoPengxiang DingSiteng HuangDonglin Wang

article2025AAAI141 citations

Introduces Cobra, a linear-complexity multimodal large language model that integrates the Mamba architecture with visual inputs to achieve faster inference speeds and comparable accuracy to LLaVA using 43% of the parameters.

Listen

Multi-modal large language models that combine visual understanding with natural language processing have gained substantial traction across vision-language reasoning, robotics, and interactive applications. However, standard models rely heavily on the Transformer architecture, which incurs a quadratic computational complexity that sharply slows processing speeds and inflates memory usage as input sequences grow. This bottleneck presents significant operational hurdles for deployment in latency-sensitive or resource-constrained settings, such as edge devices and real-time robotic feedback loops. Conventional remedies typically shrink language model capacity or heavily compress visual inputs, but these strategies often trigger steep declines in accuracy.

The article demonstrates that replacing the Transformer core with a selective state-space model allows multi-modal networks to achieve linear computational scaling and rapid inference without degrading task performance. To accomplish this, the authors develop Cobra, an architecture that couples dual pre-trained visual encoders—fusing low-level spatial features from DINOv2 with semantic features from SigLIP—to a Mamba language model backbone via a projection module. The team trained the system across roughly 1.2 million multi-turn image-text conversation samples over two epochs on eight graphics processing units, omitting the standard separate visual-text pre-alignment phase in favor of end-to-end supervised fine-tuning.

Rigorous evaluations across nine benchmark datasets spanning visual question answering, hallucination mitigation, spatial reasoning, and visual localization reveal three primary outcomes. First, Cobra achieves inference speeds between three and four times faster than leading lightweight baselines such as MobileVLM v2, reaching generation speeds of over 166 tokens per second on an enterprise graphics card. Second, a compact 3.5-billion-parameter Cobra model performs comparably to the widely used 7-billion-parameter LLaVA model while achieving higher accuracy on closed-set spatial reasoning (improving by 6.9 percentage points) and hallucination resistance (improving by 2.5 percentage points). Third, scaling the backbone to an 8-billion-parameter variant outperforms standard 7-billion-parameter baselines across all evaluated benchmarks by an average of about 6 percentage points in accuracy.

These findings indicate that state-space architectures can significantly lower deployment costs and computational latency while mitigating common model errors, such as visual hallucinations and poor spatial orientation. For enterprise systems, this architecture unlocks practical avenues for deploying high-frequency visual reasoning on edge devices and robotics platforms without requiring prohibitively expensive computing clusters. Furthermore, the results show that bypassing multi-stage pre-alignment training and avoiding aggressive visual token compression preserves vital visual fidelity without sacrificing the linear-time execution advantages of state-space models.

Organizations developing or deploying multimodal vision-language systems should consider adopting selective state-space backbones to enhance throughput and reduce operating latency. Implementers should structure prompt formats carefully, as the study shows placing optical character recognition tokens prior to queries boosts TextVQA accuracy by more than 10 percentage points due to the sequential properties of recurrent architectures. While these initial findings are strong, evaluation remains concentrated on fixed-image benchmarks using specialized server-grade graphics hardware; additional validation across dynamic video streams, diverse edge hardware profiles, and autonomous control environments will be necessary before deploying at scale.

Cover for Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Abstract

In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, current MLLMs are composed of the well-known Transformer network, which has a less efficient quadratic computation complexity. To improve the efficiency of such basic models, we propose Cobra, a linear computational complexity MLLM. Specifically, Cobra integrates the efficient Mamba language model into the visual modality. Moreover, we explore and study various modal fusion schemes to create an effective multi-modal Mamba. Extensive experiments demonstrate that (1) Cobra achieves extremely competitive performance with current computationally efficient state-of-the-art methods, e.g., LLaVA-Phi, TinyLLaVA, and MobileVLM v2, and has faster speed due to Cobra's linear sequential modeling. (2) Interestingly, the results of closed-set challenging prediction benchmarks show that Cobra performs well in overcoming visual illusions and spatial relationship judgments. (3) Notably, Cobra even achieves comparable performance to LLaVA with about 43% of the number of parameters. We will make all codes of Cobra open-source and hope that the proposed method can facilitate future research on complexity problems in MLLM. Our project page is available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 2.1 Multi-modal Large Language Models
  • 2.2 State Space Models
  • 3 Methodology
  • 3.1 Preliminaries
  • 3.2 Cobra Model
  • 3.3 Training Recipe
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Overall Performance
  • 4.3 Inference Speed
  • 4.4 Ablation Studies
  • 5 Conclusion
  • 6 Acknowledgments
  • References
  • 7 Appendix
  • 7.1 Implementation Details
  • 7.2 Additional Evaluation on TextVQA
  • 7.3 More Examples

Knowls

  1. Knowl 1 — Cobra Multi-Modal Large Language Model Architecture

    model/method

    Cobra is a multi-modal large language model (MLLM) that replaces quadratic-complexity Transformer language backbones with a linear-complexity State Space Model (Mamba) backbone to enable faster inference with constant memory usage during sequence generation.

    The Cobra architecture consists of three core components:

    1. Dual Vision Encoder: Combines DINOv2 and SigLIP ViT-SO. DINOv2 provides low-level spatial visual representations, while SigLIP provides high-level semantic features. Given an input image Xv∈RC×H×WX_v \in \mathbb{R}^{C \times H \times W}, both encoders operate on image patches to produce channel-concatenated visual tokens RvR_v.

    2. Multi-Modal Projector: A learnable Multi-Layer Perceptron (MLP) module ϕ\phi that projects concatenated visual tokens into the language model's embedding dimension DD, yielding visual prefix tokens Hv=ϕ(Rv)H_v = \phi(R_v).

    3. Mamba LLM Backbone: A sequence of Mamba selective state-space blocks (comprising short 1D convolutions, input-dependent selective SSM state transitions, RMSNorm, and residual connections). The model receives the sequence H=[Hv,Hq]∈RLin×DH = [H_v, H_q] \in \mathbb{R}^{L_{\text{in}} \times D}, where HqH_q denotes the continuous embeddings of the tokenized text instruction obtained via a GPT-NeoX tokenizer and embedding layer, and generates output tokens Y={yi}i=1LY = \{y_i\}_{i=1}^L autoregressively:

    p(Y∣Hv,Hq)=∏i=1Lp(yi∣Hv,Hq,y<i)p(Y \mid H_v, H_q) = \prod_{i=1}^L p(y_i \mid H_v, H_q, y_{<i})

  2. Knowl 2 — Dual Vision Feature Concatenation and Projection Formulation

    equation

    Given an input image Xv∈RC×H×WX_v \in \mathbb{R}^{C \times H \times W} with spatial resolution H×WH \times W and CC color channels, the image is partitioned into Nv=HWP2N_v = \frac{HW}{P^2} non-overlapping patches of size P×PP \times P.

    The dual vision encoder extracts penultimate-layer patch embeddings from DINOv2 (phiDINOv2(Xv)∈RNv×DDINOv2\\phi_{\text{DINOv2}}(X_v) \in \mathbb{R}^{N_v \times D_{\text{DINOv2}}}) and SigLIP (phiSigLIP(Xv)∈RNv×DSigLIP\\phi_{\text{SigLIP}}(X_v) \in \mathbb{R}^{N_v \times D_{\text{SigLIP}}}) and performs channel-wise concatenation to construct the unified representation Rv∈RNv×(DDINOv2+DSigLIP)R_v \in \mathbb{R}^{N_v \times (D_{\text{DINOv2}} + D_{\text{SigLIP}})}:

    Rv=[ϕDINOv2(Xv);ϕSigLIP(Xv)]R_v = [\phi_{\text{DINOv2}}(X_v); \phi_{\text{SigLIP}}(X_v)]

    The multi-modal projector ϕ:RDDINOv2+DSigLIP→RD\phi: \mathbb{R}^{D_{\text{DINOv2}} + D_{\text{SigLIP}}} \to \mathbb{R}^{D} maps the fused representations into the hidden dimension DD of the Mamba language backbone:

    Hv=ϕ(Rv)H_v = \phi(R_v)

    where Hv∈RNv×DH_v \in \mathbb{R}^{N_v \times D} is the sequence of visual prefix embeddings fed into the Mamba language model.

  3. Knowl 3 — Inference Latency and Generation Throughput Comparison

    data/table

    The generation speed and latency of Cobra models were evaluated on a single NVIDIA A100 PCIe 80GB GPU. Each model was evaluated by generating 256 output tokens in response to an image prompt with the query "Describe the image specifically". Latency TtotalT_{\text{total}} records the total wall-clock time in seconds from image encoding to the completion of 256 generated tokens, and throughput Evalavg\text{Eval}_{\text{avg}} denotes the average generation speed in tokens per second (Evalavg=256/Ttotal\text{Eval}_{\text{avg}} = 256 / T_{\text{total}}).

    Model LLM Backbone Total Params Visual Tokens Evalavg\text{Eval}_{\text{avg}} (tokens/s) TtotalT_{\text{total}} (s)
    MoE-LLaVA Phi-2-2.7B 5.3B (3.6B Act.) 576 20.33 12.59
    LLaVA-Phi Phi-2-2.7B 3.1B 576 40.89 6.26
    MobileVLM v2 MobileLLaMA-2.7B 3.1B 144 49.50 5.17
    Cobra-3.5B (ours) Mamba-2.8B 3.5B 729 166.47 1.54
    Cobra-LDPv2-3.5B (ours) Mamba-2.8B 3.5B 196 166.85 1.53
    Cobra-8B (ours) Mamba-7B 7.8B 729 79.92 3.20

    Cobra-3.5B achieves 166.47 tokens/s (1.54 s total), which is approximately 3.36×3.36\times to 4.07×4.07\times faster than lightweight Transformer baselines (MobileVLM v2 at 49.50 tokens/s and LLaVA-Phi at 40.89 tokens/s) and >8×>8\times faster than MoE-LLaVA (20.33 tokens/s), despite processing a higher number of visual tokens (729 vs. 144/576). Compressing visual tokens to 196 via an LDPv2 downsampling projector yielded negligible speedup (166.85 tokens/s) because in recurrent SSM architectures, prompt length only impacts the initial parallel forward pass rather than autoregressive token generation.

  4. Knowl 4 — Multi-Modal Benchmark Performance Across Model Scales

    data/table

    Evaluation of Cobra and baseline multi-modal LLMs across four open-ended visual question answering benchmarks (VQA-v2, GQA, VizWiz, TextVQA) and two closed-set benchmarks evaluating spatial reasoning (VSR) and object hallucination (POPE). Res. indicates input image resolution.

    Model LLM Backbone Res. VQA-v2 GQA VizWiz TextVQA VSR POPE
    Large Scale MLLMs
    OpenFlamingo MPT-7B 336 52.7 – 27.5 33.6 – –
    BLIP-2 Vicuna-13B 224 – 41.0 19.6 42.5 50.9 –
    MiniGPT-4 Vicuna-7B 224 32.2 – – – – –
    InstructBLIP Vicuna-7B 224 – 49.2 34.5 50.1 54.3 –
    InstructBLIP Vicuna-13B 224 – 49.5 33.4 50.7 52.1 –
    Shikra Vicuna-13B 224 77.4 – – – – –
    IDEFICS LLaMA-7B 224 50.9 – 35.5 25.9 – –
    IDEFICS LLaMA-75B 224 60.0 – 36.0 30.9 – –
    Qwen-VL Qwen-7B 448 78.2 59.3 35.2 63.8 – –
    LLaVA v1.5 Vicuna-7B 336 78.5 62.0 50.0 58.2 51.5 85.9
    Cobra-8B (ours) Mamba-7B 384 79.2 63.9 56.2 59.5 62.9 87.6
    Small Scale MLLMs
    MoE-LLaVA StableLM-1.6B 336 76.7 60.3 36.2 50.1 – 85.7
    MoE-LLaVA Phi2-2.7B 384 79.9 62.6 43.7 57.0 – 85.7
    LLaVA-Phi Phi2-2.7B 336 71.4 – 35.9 48.6 – 85.0
    MobileVLM v2 MobileLLaMA-2.7B 336 – 61.1 – 57.5 – 84.7
    Cobra-3.5B (ours) Mamba-2.8B 384 77.8 62.3 49.7 58.2 58.4 88.4

    Cobra-8B attains top performance across all evaluated datasets, improving over LLaVA v1.5-7B by +11.4% on VSR (62.9% vs. 51.5%), +6.2% on VizWiz (56.2% vs. 50.0%), and +1.7% on POPE (87.6% vs. 85.9%). Cobra-3.5B performs competitively with LLaVA v1.5-7B using ~48% of its parameters, scoring +6.9% higher on VSR (58.4% vs. 51.5%) and +2.5% higher on POPE (88.4% vs. 85.9%).

  5. Knowl 5 — Direct End-to-End Fine-Tuning Training Scheme for Cobra

    model/method

    Instead of the conventional two-stage LLaVA training recipe (which first trains only the projection layer during a pre-alignment stage before fine-tuning the full LLM for one epoch), Cobra eliminates the isolated pre-alignment phase and trains the projector and the entire Mamba LLM backbone jointly end-to-end for 2 epochs.

    The training dataset contains approximately 1.2 million images with associated dialogue instructions combining:

    1. The LLaVA v1.5 multi-modal mixture (655K samples, covering academic VQA, LLaVA-Instruct visual instruction tuning, and ShareGPT pure-text dialogue).
    2. LVIS-Instruct-4V (220K images with GPT-4V generated context-aware instructions).
    3. LRV-Instruct (400K visual instruction instances spanning 16 vision-language tasks targeting hallucination reduction).

    Hyperparameters and optimization settings include:

    • Image resolution: 384×384384 \times 384 (Nv=729N_v = 729 visual tokens)
    • Optimizer: AdamW with weight decay 0.10.1
    • Learning rate schedule: Cosine decay with initial learning rate 2×10−52 \times 10^{-5} and warm-up ratio 0.030.03
    • Global batch size: 128 across 8 NVIDIA A100 80GB GPUs using PyTorch Fully Sharded Data Parallel (FSDP) with FP32/BF16 mixed precision
    • Total training steps: 19K steps (~26.5 hours for Cobra-3.5B)
  6. Knowl 6 — Prompt Order Sensitivity in Recurrent State-Space MLLMs

    empirical result

    In Mamba-based MLLMs, the relative order of text prompt components has a significant impact on downstream task performance due to the sequential/recurrent inductive bias of state space models.

    On the TextVQA benchmark, feeding Optical Character Recognition (OCR) tokens following the standard LLaVA prompt order ("OCR Last": Question\n Reference OCR token: ...) caused a severe drop in Cobra-3.5B accuracy to 43.0%, which is lower than providing no OCR tokens at all (47.9%). Inverting the prompt format to place the OCR tokens first ("OCR First": Reference OCR token: ...\n Question) resolved the issue, raising accuracy to 58.2% (a +15.2% absolute improvement over "OCR Last").

  7. Knowl 7 — Ablation Analysis of Vision Encoders, Projectors, and Training Schemes

    data/table

    Ablation results on Cobra-3.5B evaluating the contribution of the dual vision encoder (concatenated DINOv2 + SigLIP vs. SigLIP alone), projection module (standard MLP vs. LDPv2 lightweight downsampling projector), language model initialization (instruction-tuned Mamba-2.8b-Zephyr vs. pre-trained Base Mamba-2.8B), and training duration/scheme (2-epoch direct fine-tuning vs. 1-epoch fine-tuning vs. pre-training followed by fine-tuning PT+FT).

    Model Variant VQA-v2 GQA VizWiz TextVQA VSR POPE RefCOCO RefCOCO+ RefCOCOg
    Cobra-3.5B (Full) 77.8 62.3 49.7 58.2 58.4 88.4 52.7 45.6 48.9
    w/ SigLIP 77.5 61.8 48.3 58.8 53.2 88.2 46.7 40.1 43.8
    w/ LDPv2 76.2 61.9 50.2 54.7 56.1 87.7 50.3 42.9 46.9
    w/ Base 77.8 62.7 47.2 57.9 54.4 89.0 52.2 45.6 48.6
    w/ 1 Ep FT 76.5 60.9 48.5 57.5 53.8 88.1 42.5 34.3 39.0
    w/ PT+FT 75.7 60.4 44.2 58.0 51.6 86.9 37.3 29.7 34.3

    Key empirical findings from the ablation study include:

    1. Removing DINOv2 (w/ SigLIP) degrades spatial reasoning (VSR drops from 58.4 to 53.2) and visual grounding (RefCOCO drops by 5.1 to 6.0 points).
    2. Using a downsampling projector (w/ LDPv2) hurts tasks requiring precise local details (TextVQA drops by 3.5, VSR by 2.3).
    3. Single-epoch fine-tuning (w/ 1 Ep FT) underfits, severely degrading grounding benchmarks by ~10 points.
    4. Initializing with a pre-aligned projector before fine-tuning (w/ PT+FT) performs worse across all benchmarks than training end-to-end directly.
  8. Knowl 8 — Visual Grounding Benchmark Evaluation

    data/table

    Performance of Cobra variants compared with LLaVA v1.5-7B across visual grounding benchmarks (RefCOCO, RefCOCO+, RefCOCOg). Grounding accuracy measures the model's ability to locate image regions from natural language referring expressions of varying lengths and styles.

    Model RefCOCO RefCOCO+ RefCOCOg Avg.
    LLaVA v1.5 (Vicuna-7B) 55.1 49.5 50.9 51.8
    Cobra-3.5B (Mamba-2.8B) 52.7 45.6 46.9 48.4
    Cobra-8B (Mamba-7B) 58.2 52.5 54.4 55.0

    Cobra-8B outperforms LLaVA v1.5-7B by +3.2% on average (55.0% vs. 51.8%), exceeding it on RefCOCO (58.2% vs. 55.1%), RefCOCO+ (52.5% vs. 49.5%), and RefCOCOg (54.4% vs. 50.9%). Cobra-3.5B trails LLaVA v1.5-7B by 3.4% on average (48.4% vs. 51.8%), indicating that visual referring capability scales directly with the capacity of the underlying language model.

  9. Knowl 9 — Impact of OCR Prompt Formatting on TextVQA Performance

    data/table

    TextVQA accuracy under three prompt variations: OCR First (placing OCR token context before the question), OCR Last (placing the question before OCR tokens), and w/o OCR tokens (providing no OCR context).

    Model OCR First OCR Last w/o OCR tokens
    LLaVA v1.5 – 58.2 46.1
    Cobra-3.5B 58.2 43.0 47.9
    w/ SigLIP 58.8 47.3 49.3
    w/ LDPv2 54.7 44.7 40.3
    w/ Base 57.9 47.6 47.9
    w/ 1 Ep FT 57.5 45.4 46.4
    w/ PT+FT 58.0 47.4 46.6
    Cobra-8B 59.5 43.0 50.7

    Across all Cobra configurations and model sizes, placing OCR tokens first (OCR First) yields a consistent performance advantage of over 10% compared to OCR Last. In OCR Last, the presence of OCR context actually lowers accuracy below the zero-context baseline (w/o OCR tokens) for most models (e.g., Cobra-3.5B drops from 47.9% to 43.0%, Cobra-8B drops from 50.7% to 43.0%).

Coverage note — Qualitative dialogue and visual captioning visualization examples (Tables 6, 8, 9, 10) were omitted as they serve purely illustrative purposes without conveying standalone quantitative findings.

References

  1. 1.Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; et al. 2022. Flamingo: a Visual Language Model for Few-Shot Learning. arXiv:2204.14198.
  2. 2.Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, 2425–2433.
  3. 3.Awadalla, A.; Gao, I.; Gardner, J.; Hessel, J.; Hanafy, Y.; et al. 2023. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models. arXiv:2308.01390.
  4. 4.Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; et al. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966.
  5. 5.Bar-David, S.; Zimerman, I.; Nachmani, E.; and Wolf, L. 2023. Decision S4: Efficient Sequence-Based RL via State Spaces Layers. arXiv:2306.05167.
  6. 6.Bellagente, M.; Tow, J.; Mahan, D.; Phung, D.; Zhuravinskyi, M.; Adithyan, R.; Baicoianu, J.; Brooks, B.; Cooper, N.; Datta, A.; Lee, M.; Mostaque, E.; Pieler, M.; Pinnaparju, N.; Rocha, P.; Saini, H.; Teufel, H.; Zanichelli, N.; and Riquelme, C. 2024. Stable LM 2 1.6B Technical Report. arXiv:2402.17834.
  7. 7.Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; et al. 2023. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818.
  8. 8.Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; et al. 2023a. Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic. arXiv:2306.15195.
  9. 9.Chen, L.; Li, J.; Dong, X.; Zhang, P.; He, C.; Wang, J.; et al. 2023b. ShareGPT4V: Improving Large Multi-Modal Models with Better Captions. arXiv:2311.12793.
  10. 10.Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality.
  11. 11.Chu, X.; Qiao, L.; Lin, X.; Xu, S.; Yang, Y.; Hu, Y.; et al. 2023. MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices. arXiv:2312.16886.
  12. 12.Chu, X.; Qiao, L.; Zhang, X.; Xu, S.; Wei, F.; Yang, Y.; et al. 2024. MobileVLM V2: Faster and Stronger Baseline for Vision Language Model. arXiv:2402.03766.
  13. 13.Cui, G.; Yuan, L.; Ding, N.; Yao, G.; Zhu, W.; Ni, Y.; et al. 2023. UltraFeedback: Boosting Language Models with High-quality Feedback. arXiv:2310.01377.
  14. 14.Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; et al. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv:2305.06500.
  15. 15.Dao, T.; and Gu, A. 2024. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. arXiv:2405.21060.
  16. 16.Ding, N.; Chen, Y.; Xu, B.; Qin, Y.; Zheng, Z.; Hu, S.; et al. 2023. Enhancing Chat Language Models by Scaling Highquality Instructional Conversations. arXiv:2305.14233.
  17. 17.Ding, P.; Zhao, H.; Song, W.; Zhang, W.; Zhang, M.; et al. 2024. QUAR-VLA: Vision-Language-Action Model for Quadruped Robots. arXiv:2312.14457.
  18. 18.Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; et al. 2023. LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model. arXiv:2304.15010.
  19. 19.Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. arXiv:1612.00837.
  20. 20.Gu, A.; and Dao, T. 2023. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752.
  21. 21.Gu, A.; Goel, K.; and Re, C. 2022. Efficiently Modeling Long Sequences with Structured State Spaces. arXiv:2111.00396.
  22. 22.Gurari, D.; Li, Q.; Stangl, A. J.; Guo, A.; Lin, C.; et al. 2018. VizWiz Grand Challenge: Answering Visual Questions from Blind People. arXiv:1802.08218.
  23. 23.Hasani, R.; Lechner, M.; Wang, T.-H.; Chahine, M.; Amini, A.; and Rus, D. 2022. Liquid Structural State-Space Models. arXiv:2209.12951.
  24. 24.Hendriksen, M.; Vakulenko, S.; Kuiper, E.; and de Rijke, M. 2023. Scene-centric vs. object-centric image-text crossmodal retrieval: a reproducibility study. In European Conference on Information Retrieval, 68–85. Springer.
  25. 25.Hudson, D. A.; and Manning, C. D. 2019. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. arXiv:1902.09506.
  26. 26.Karamcheti, S.; Nair, S.; Balakrishna, A.; Liang, P.; Kollar, T.; et al. 2024. Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models. arXiv:2402.07865.
  27. 27.Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. arXiv:2006.16236.
  28. 28.Kazemzadeh, S.; Ordonez, V.; andre Matten, M.; and Berg, T. L. 2014. ReferItGame: Referring to Objects in Photographs of Natural Scenes. In Conference on Empirical Methods in Natural Language Processing.
  29. 29.Ke, L.; Pei, W.; Li, R.; Shen, X.; and Tai, Y.-W. 2019. Reflective decoding network for image captioning. In Proceedings of the IEEE/CVF international conference on computer vision, 8888–8897.
  30. 30.Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; et al. 2024. OpenVLA: An Open-Source VisionLanguage-Action Model. arXiv:2406.09246.
  31. 31.Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; et al. 2016. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. arXiv:1602.07332.
  32. 32.Laurenc¸on, H.; Saulnier, L.; Tronchon, L.; Bekman, S.; Singh, A.; et al. 2023. OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents. arXiv:2306.16527.
  33. 33.Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023a. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. arXiv:2301.12597.
  34. 34.Li, Y.; Du, Y.; Zhou, K.; Wang, J.; Zhao, W. X.; et al. 2023b. Evaluating Object Hallucination in Large Vision-Language Models. arXiv:2305.10355.
  35. 35.Lin, B.; Tang, Z.; Ye, Y.; Cui, J.; Zhu, B.; Jin, P.; et al. 2024. MoE-LLaVA: Mixture of Experts for Large VisionLanguage Models. arXiv:2401.15947.
  36. 36.Liu, F.; Emerson, G.; and Collier, N. 2023. Visual Spatial Reasoning. arXiv:2205.00363.
  37. 37.Liu, F.; Lin, K.; Li, L.; Wang, J.; Yacoob, Y.; et al. 2023a. Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. arXiv:2306.14565.
  38. 38.Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2023b. Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744.
  39. 39.Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023c. Visual Instruction Tuning. arXiv:2304.08485.
  40. 40.Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. arXiv:1711.05101.
  41. 41.Lu, C.; Schroecker, Y.; Gu, A.; Parisotto, E.; Foerster, J.; et al. 2023. Structured State Space Models for In-Context Reinforcement Learning. arXiv:2303.03982.
  42. 42.Lu, J.; Clark, C.; Zellers, R.; Mottaghi, R.; and Kembhavi, A. 2022. Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks. arXiv:2206.08916.
  43. 43.Mercat, J.; Vasiljevic, I.; Keh, S.; Arora, K.; Dave, A.; Gaidon, A.; and Kollar, T. 2024. Linearizing Large Language Models. arXiv:2405.06640.
  44. 44.OpenAI. 2023. GPT-4V(ision) System Card.
  45. 45.Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; et al. 2024. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193.
  46. 46.Penedo, G.; Malartic, Q.; Hesslow, D.; Cojocaru, R.; Cappelli, A.; et al. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116.
  47. 47.Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290.
  48. 48.ShareGPT. 2023. https://sharegpt.com/.
  49. 49.Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; et al. 2019. Towards VQA Models That Can Read. arXiv:1904.08920.
  50. 50.Smith, J. T. H.; Warrington, A.; and Linderman, S. W. 2023. Simplified State Space Layers for Sequence Modeling. arXiv:2208.04933.
  51. 51.Soboleva, D.; Al-Khateeb, F.; Myers, R.; Steeves, J. R.; Hestness, J.; and Dey, N. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https://www.cerebras.net/blog/slimpajama-a-627btoken-cleaned-and-deduplicated-version-of-redpajama.
  52. 52.Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023. Stanford Alpaca: An Instruction-following LLaMA model. https://github.com/tatsu-lab/stanford_alpaca.
  53. 53.Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. arXiv:2401.06209.
  54. 54.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
  55. 55.Waleffe, R.; Byeon, W.; Riach, D.; Norick, B.; Korthikanti, V.; et al. 2024. An Empirical Study of Mamba-based Language Models. arXiv:2406.07887.
  56. 56.Wang, J.; Meng, L.; Weng, Z.; He, B.; Wu, Z.; et al. 2023. To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning. arXiv:2311.07574.
  57. 57.Wightman, R. 2019. PyTorch Image Models. https://github.com/rwightman/pytorch-image-models.
  58. 58.Wu, S.; Fei, H.; Qu, L.; Ji, W.; and Chua, T.-S. 2024. NExTGPT: Any-to-Any Multimodal LLM. arXiv:2309.05519.
  59. 59.Yan, J. N.; Gu, J.; and Rush, A. M. 2023. Diffusion Models Without Attention. arXiv:2311.18257.
  60. 60.Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016. Modeling Context in Referring Expressions. arXiv:1608.00272.
  61. 61.Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023. Sigmoid Loss for Language Image Pre-Training. arXiv:2303.15343.
  62. 62.Zhang, B.; and Sennrich, R. 2019. Root Mean Square Layer Normalization. arXiv:1910.07467.
  63. 63.Zhang, P.; Zeng, G.; Wang, T.; and Lu, W. 2024. TinyLlama: An Open-Source Small Language Model. arXiv:2401.02385.
  64. 64.Zhao, Y.; Gu, A.; Varma, R.; Luo, L.; Huang, C.-C.; et al. 2023. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel. arXiv:2304.11277.
  65. 65.Zhou, B.; Hu, Y.; Weng, X.; Jia, J.; Luo, J.; Liu, X.; et al. 2024. TinyLLaVA: A Framework of Small-scale Large Multimodal Models. arXiv:2402.14289.
  66. 66.Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv:2304.10592.
  67. 67.Zhu, Y.; Zhu, M.; Liu, N.; Ou, Z.; Mou, X.; and Tang, J. 2024. LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model. arXiv:2401.02330.

Citation

MLA
Zhao, H., et al. “Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 10, 2025, pp. 10421–29, https://doi.org/10.1609/aaai.v39i10.33131.
APA
Zhao, H., Zhang, M., Zhao, W., Ding, P., Huang, S., & Wang, D. (2025). Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference. Proceedings of the AAAI Conference on Artificial Intelligence, 39(10), 10421–10429. https://doi.org/10.1609/aaai.v39i10.33131
Chicago
Zhao, H., M. Zhang, W. Zhao, P. Ding, S. Huang, and D. Wang. 2025. “Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference”. Proceedings of the AAAI Conference on Artificial Intelligence 39 (10): 10421–29. https://doi.org/10.1609/aaai.v39i10.33131.
Harvard
Zhao, H. et al. (2025) “Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference”, Proceedings of the AAAI Conference on Artificial Intelligence, 39(10), pp. 10421–10429. Available at: https://doi.org/10.1609/aaai.v39i10.33131.
Vancouver
1. Zhao H, Zhang M, Zhao W, Ding P, Huang S, Wang D (2025) Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference. Proceedings of the AAAI Conference on Artificial Intelligence 39:10421–10429

BibTeX

@article{Zhao_2025, title={Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference}, volume={39}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v39i10.33131}, DOI={10.1609/aaai.v39i10.33131}, number={10}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhao, Han and Zhang, Min and Zhao, Wei and Ding, Pengxiang and Huang, Siteng and Wang, Donglin}, year={2025}, month=Apr, pages={10421–10429} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF