Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
Han ZhaoMin ZhangWei ZhaoPengxiang DingSiteng HuangDonglin Wang
Introduces Cobra, a linear-complexity multimodal large language model that integrates the Mamba architecture with visual inputs to achieve faster inference speeds and comparable accuracy to LLaVA using 43% of the parameters.
Multi-modal large language models that combine visual understanding with natural language processing have gained substantial traction across vision-language reasoning, robotics, and interactive applications. However, standard models rely heavily on the Transformer architecture, which incurs a quadratic computational complexity that sharply slows processing speeds and inflates memory usage as input sequences grow. This bottleneck presents significant operational hurdles for deployment in latency-sensitive or resource-constrained settings, such as edge devices and real-time robotic feedback loops. Conventional remedies typically shrink language model capacity or heavily compress visual inputs, but these strategies often trigger steep declines in accuracy.
The article demonstrates that replacing the Transformer core with a selective state-space model allows multi-modal networks to achieve linear computational scaling and rapid inference without degrading task performance. To accomplish this, the authors develop Cobra, an architecture that couples dual pre-trained visual encoders—fusing low-level spatial features from DINOv2 with semantic features from SigLIP—to a Mamba language model backbone via a projection module. The team trained the system across roughly 1.2 million multi-turn image-text conversation samples over two epochs on eight graphics processing units, omitting the standard separate visual-text pre-alignment phase in favor of end-to-end supervised fine-tuning.
Rigorous evaluations across nine benchmark datasets spanning visual question answering, hallucination mitigation, spatial reasoning, and visual localization reveal three primary outcomes. First, Cobra achieves inference speeds between three and four times faster than leading lightweight baselines such as MobileVLM v2, reaching generation speeds of over 166 tokens per second on an enterprise graphics card. Second, a compact 3.5-billion-parameter Cobra model performs comparably to the widely used 7-billion-parameter LLaVA model while achieving higher accuracy on closed-set spatial reasoning (improving by 6.9 percentage points) and hallucination resistance (improving by 2.5 percentage points). Third, scaling the backbone to an 8-billion-parameter variant outperforms standard 7-billion-parameter baselines across all evaluated benchmarks by an average of about 6 percentage points in accuracy.
These findings indicate that state-space architectures can significantly lower deployment costs and computational latency while mitigating common model errors, such as visual hallucinations and poor spatial orientation. For enterprise systems, this architecture unlocks practical avenues for deploying high-frequency visual reasoning on edge devices and robotics platforms without requiring prohibitively expensive computing clusters. Furthermore, the results show that bypassing multi-stage pre-alignment training and avoiding aggressive visual token compression preserves vital visual fidelity without sacrificing the linear-time execution advantages of state-space models.
Organizations developing or deploying multimodal vision-language systems should consider adopting selective state-space backbones to enhance throughput and reduce operating latency. Implementers should structure prompt formats carefully, as the study shows placing optical character recognition tokens prior to queries boosts TextVQA accuracy by more than 10 percentage points due to the sequential properties of recurrent architectures. While these initial findings are strong, evaluation remains concentrated on fixed-image benchmarks using specialized server-grade graphics hardware; additional validation across dynamic video streams, diverse edge hardware profiles, and autonomous control environments will be necessary before deploying at scale.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Introduces the selective state space model Mamba, whose linear-time sequential modeling Cobra adapts as its core language backbone for efficient multimodal inference.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Establishes the standard LLaVA multimodal architecture and visual instruction tuning recipe that Cobra uses as its primary baseline and structural template.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). Pioneers the application of Mamba's selective state space mechanisms to visual representation learning, providing foundational context for linear-time vision processing.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). Develops a 2D selective scan mechanism to adapt Mamba for visual backbones, directly informing Cobra's exploration of state space models across visual modalities.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). Investigates the architectural design space and visual representation fusion schemes that Cobra builds upon when integrating vision features with state space models.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Surveys the core architectural pipelines, connector interfaces, and training strategies used across multimodal large language models.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). Introduces the comprehensive MME benchmark utilized by Cobra to evaluate perception, cognition, and visual illusion metrics.
- Paper: Mamba-3: Improved Sequence Modeling using State Space Principles, Aakash Lahoti et al. (2026). Extends state space sequence modeling principles into Mamba-3, offering advanced SSM formulations that can further improve upon Cobra's linear-time multimodal inference.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). Develops complementary training-free token reduction methods to accelerate multimodal inference latency beyond linear architectural backbones like Cobra.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Addresses high-resolution visual token bottlenecks with dynamic compression, offering an orthogonal avenue for inference acceleration alongside linear-complexity MLLMs.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Surveys recent efficient sequence modeling and multimodal architectures, contextualizing state space models like Cobra within broader non-quadratic LLM paradigms.
