MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Ali HatamizadehJan Kautz

article2025CVPR533 citations

Introduces a hybrid vision backbone that strategically places self-attention layers after redesigned Mamba blocks, establishing a new Pareto frontier for the trade-off between ImageNet-1K accuracy and image throughput.

Listen

Modern computer vision systems rely heavily on Transformer models due to their exceptional ability to capture complex spatial context across full images. However, Transformers suffer from severe computational bottlenecks because their processing cost scales quadratically with image resolution, making them expensive and slow to deploy in production. While recent state-space sequence models like Mamba offer linear-time computational efficiency, their directional, step-by-step design fundamentally struggles to process visual data, where spatial context must be evaluated holistically rather than in a single sequence. Previous adaptations attempted to bridge this gap using complex bidirectional scans, but these added significant processing latency and memory overhead.

The article demonstrates a novel hybrid architecture, named MambaVision, designed specifically for visual recognition tasks. The main objective was to develop and evaluate an architecture that integrates redesigned state-space models with traditional self-attention mechanisms, capturing long-range spatial context while maximizing image throughput and minimizing computational requirements.

The researchers conducted comprehensive empirical evaluations using standard computer vision benchmarks. They evaluated image classification on the ImageNet-1K and ImageNet-21K datasets, as well as downstream tasks including object detection and instance segmentation on MS COCO and semantic segmentation on ADE20K. The architecture uses a four-stage hierarchical approach: the first two stages apply standard convolutional layers for fast initial feature extraction, while the final two stages integrate redesigned Mamba blocks alongside standard Transformer self-attention blocks.

The key findings demonstrate major performance and efficiency advantages. First, MambaVision establishes a superior trade-off between accuracy and image throughput across all tested scales. For example, the base MambaVision model achieves 84.2% top-1 accuracy on ImageNet-1K while processing 3,670 images per second, outperforming comparable models like ConvNeXt-B (83.8% at 1,485 images per second) and VMamba-B (83.9% at 645 images per second). Second, the architecture reduces computational load significantly, with the base model requiring 56% fewer floating-point operations than comparable high-end vision models. Third, MambaVision consistently outperforms competing backbones on downstream tasks, showing higher average precision in object detection and instance segmentation on MS COCO and higher intersection-over-union scores on ADE20K. Finally, ablation studies confirm that placing Transformer self-attention blocks specifically in the final layers of the deep stages is essential for recovering global context and achieving peak accuracy.

These findings indicate that organizations can achieve state-of-the-art visual accuracy while dramatically reducing the hardware latency and cloud computing costs associated with pure Transformer models. The results challenge the assumption that pure state-space models can entirely replace self-attention in visual domains, showing instead that a deliberate hybrid combination provides the optimal balance of speed and contextual understanding.

Organizations developing or deploying high-throughput computer vision applications should consider adopting hybrid architectures like MambaVision as efficient backbone models. For operational deployment, engineering teams should benchmark MambaVision variants directly against their existing convolutional or Transformer models to quantify latency and cost reductions on target hardware.

The primary limitation of the study is that evaluations were conducted in standard academic benchmark settings using high-end server GPUs, without covering ultra-low-power edge devices or custom domain datasets. Nonetheless, confidence in the reported performance is high given the rigorous cross-benchmark validation and open-source availability of the implementation.

arXiv: 2407.08083NVlabs/MambaVision
Cover for MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Abstract

We propose a novel hybrid Mamba-Transformer backbone, MambaVision, specifically tailored for vision applications. Our core contribution includes redesigning the Mamba formulation to enhance its capability for efficient modeling of visual features. Through a comprehensive ablation study, we demonstrate the feasibility of integrating Vision Transformers (ViT) with Mamba. Our results show that equipping the Mamba architecture with self-attention blocks in the final layers greatly improves its capacity to capture long-range spatial dependencies. Based on these findings, we introduce a family of MambaVision models with a hierarchical architecture to meet various design criteria. For classification on the ImageNet-1K dataset, MambaVision variants achieve state-of-the-art (SOTA) performance in terms of both Top-1 accuracy and throughput. In downstream tasks such as object detection, instance segmentation, and semantic segmentation on MS COCO and ADE20K datasets, MambaVision outperforms comparably sized backbones while demonstrating favorable performance. Code: https://github.com/NVlabs/MambaVision

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Macro Architecture
  • 3.2. Micro Architecture
  • 3.2.1. Mamba Preliminaries
  • 3.2.2. Layer Architecture
  • 4. Experiments
  • 5. Results
  • 5.1. Image classification
  • 5.2. Object Detection and Segmentation
  • 5.3. Ablation
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — MambaVision Macro Architecture

    model/method

    MambaVision is a hierarchical four-stage vision backbone designed to balance computational throughput and global context modeling.

    Given an input image of dimensions H×W×3H \times W \times 3, the macro architecture processes representations across four stages:

    1. Patch Embedding Stem: Two consecutive 3×33 \times 3 2D convolutional layers with stride 2 project the image into a CC-dimensional feature space of size H4×W4×C\frac{H}{4} \times \frac{W}{4} \times C.

    2. Downsamplers: Between consecutive stages, downsampling is performed using a 3×33 \times 3 2D convolutional layer with stride 2, reducing the spatial resolution by half and doubling the channel dimension.

    3. Stages 1 and 2 (CNN Residual Stages): Early stages operate on high-resolution feature maps using standard residual convolutional blocks to maximize feature extraction speed: z^=GELU(BN(Conv3×3(z)))\hat{z} = \text{GELU}(\text{BN}(\text{Conv}_{3\times3}(z))) z=BN(Conv3×3(z^))+zz = \text{BN}(\text{Conv}_{3\times3}(\hat{z})) + z where Conv3×3\text{Conv}_{3\times3} denotes a 3×33 \times 3 convolution, BN\text{BN} is Batch Normalization, and GELU\text{GELU} is the Gaussian Error Linear Unit activation.

    4. Stages 3 and 4 (Hybrid Stages): Lower-resolution stages employ a hybrid design. For a stage containing NN total layers, the first N2\frac{N}{2} layers consist of MambaVision mixer blocks followed by Multi-Layer Perceptron (MLP) blocks, while the remaining N2\frac{N}{2} layers consist of Multi-Head Self-Attention blocks followed by MLP blocks.

    5. Classification Head: Feature maps from Stage 4 undergo 2D Adaptive Average Pooling followed by a Linear projection layer to generate class logits.

  2. Knowl 2 — MambaVision Token Mixer Formulation

    model/method

    The MambaVision token mixer redesigns the standard State Space Model (SSM) block to better handle 2D visual representations by removing causal constraints and introducing a symmetric non-SSM branch.

    For an input token representation Xin∈RT×CX_{\text{in}} \in \mathbb{R}^{T \times C}, where TT is the sequence length and CC is the channel embedding dimension, the mixer branches the computation into two parallel paths, each mapped to channel dimension C2\frac{C}{2}:

    X1=Scan(σ(Conv1D(Linear(C,C/2)(Xin))))X_1 = \text{Scan}(\sigma(\text{Conv}_{1\text{D}}(\text{Linear}(C, C/2)(X_{\text{in}})))) X2=σ(Conv1D(Linear(C,C/2)(Xin)))X_2 = \sigma(\text{Conv}_{1\text{D}}(\text{Linear}(C, C/2)(X_{\text{in}}))) Xout=Linear(C,C)(Concat(X1,X2))X_{\text{out}} = \text{Linear}(C, C)(\text{Concat}(X_1, X_2))

    Here:

    • Linear(Cin,Cout)\text{Linear}(C_{\text{in}}, C_{\text{out}}) is a linear projection layer from dimension CinC_{\text{in}} to CoutC_{\text{out}}.
    • Conv1D\text{Conv}_{1\text{D}} denotes a non-causal regular 1D convolution with kernel size 3 and 'same' padding, applied independently to each branch.
    • σ\sigma is the Sigmoid Linear Unit (SiLU) activation function.
    • Scan(⋅)\text{Scan}(\cdot) denotes the input-dependent selective state space scan operation with discretized parameters (Aˉ,Bˉ,Cˉ,D)(\bar{A}, \bar{B}, \bar{C}, D).
    • Concat(X1,X2)∈RT×C\text{Concat}(X_1, X_2) \in \mathbb{R}^{T \times C} concatenates the outputs of the SSM branch X1∈RT×C/2X_1 \in \mathbb{R}^{T \times C/2} and the non-SSM spatial branch X2∈RT×C/2X_2 \in \mathbb{R}^{T \times C/2} along the channel dimension.
    • The final Linear(C,C)\text{Linear}(C, C) projects the concatenated representation back to dimension CC.
  3. Knowl 3 — PyTorch Implementation of MambaVision Mixer

    algorithm

    The MambaVision token mixer module is implemented in PyTorch as follows:

    import math
    import torch
    import torch.nn as nn
    import torch.nn.functional as F
    from einops import rearrange, repeat
    
    class MambaVisionMixer(nn.Module):
        def __init__(self, dim, d_state=16, kernel_size=3, dt_min=0.001, dt_max=0.1):
            super().__init__()
            self.dim = dim
            self.d_state = d_state
            self.dt_rank = math.ceil(dim / 16)
            self.in_proj = nn.Linear(dim, dim)
            self.x_proj = nn.Linear(dim // 2, self.dt_rank + self.d_state * 2)
            self.conv1d_x = nn.Conv1d(
                dim // 2, dim // 2, kernel_size=kernel_size, padding='same', groups=dim // 2
            )
            self.conv1d_z = nn.Conv1d(
                dim // 2, dim // 2, kernel_size=kernel_size, padding='same', groups=dim // 2
            )
            self.dt_proj = nn.Linear(self.dt_rank, dim // 2)
            dt = torch.exp(
                torch.rand(dim // 2) * (math.log(dt_max) - math.log(dt_min)) + math.log(dt_min)
            )
            A_log = torch.log(repeat(torch.arange(1, self.d_state + 1), 'n -> d n', d=dim // 2))
            self.A_log = nn.Parameter(A_log)
            self.D = nn.Parameter(torch.ones(dim // 2))
            self.out_proj = nn.Linear(dim, dim)
    
        def forward(self, hidden_states):
            # hidden_states: (b, l, d)
            xz = rearrange(self.in_proj(hidden_states), 'b l d -> b d l')
            x, z = xz.chunk(2, dim=1)
            A = -torch.exp(self.A_log)
            x = F.silu(self.conv1d_x(x))
            z = F.silu(self.conv1d_z(z))
            seqlen = hidden_states.shape[1]
            x_dbl = self.x_proj(rearrange(x, 'b d l -> (b l) d'))
            dt, B, C = torch.split(x_dbl, [self.dt_rank, self.d_state, self.d_state], dim=-1)
            dt = rearrange(self.dt_proj(dt), '(b l) d -> b d l', l=seqlen)
            B = rearrange(B, '(b l) dstate -> b dstate l', l=seqlen)
            C = rearrange(C, '(b l) dstate -> b dstate l', l=seqlen)
            x_ssm = selective_scan_fn(x, dt, A, B, C, self.D)
            hidden_states = rearrange(torch.cat([x_ssm, z], dim=1), 'b d l -> b l d')
            return self.out_proj(hidden_states)
    
  4. Knowl 4 — Hybrid Layer Architecture and Self-Attention Integration

    model/method

    In stages 3 and 4 of MambaVision, each stage contains NN consecutive layers. Given an input token tensor Xn−1∈RT×CX^{n-1} \in \mathbb{R}^{T \times C} at layer n∈{1,…,N}n \in \{1, \dots, N\}, the layer computation follows:

    X^n=Mixer(Norm(Xn−1))+Xn−1\hat{X}^n = \text{Mixer}(\text{Norm}(X^{n-1})) + X^{n-1} Xn=MLP(Norm(X^n))+X^nX^n = \text{MLP}(\text{Norm}(\hat{X}^n)) + \hat{X}^n

    where Norm\text{Norm} denotes Layer Normalization, MLP\text{MLP} denotes a standard two-layer feedforward network with GELU activations, and Mixer\text{Mixer} is selected based on layer index nn:

    • For the first N2\frac{N}{2} layers (n∈{1,…,N2}n \in \{1, \dots, \frac{N}{2}\}), Mixer\text{Mixer} is the MambaVision token mixer.
    • For the remaining N2\frac{N}{2} layers (n∈{N2+1,…,N}n \in \{\frac{N}{2} + 1, \dots, N\}), Mixer\text{Mixer} is Multi-Head Self-Attention (MHSA): Attention(Q,K,V)=Softmax(QK⊤dh)V\text{Attention}(Q, K, V) = \text{Softmax}\left(\frac{QK^\top}{\sqrt{d_h}}\right)V where Q,K,V∈RTw×dhQ, K, V \in \mathbb{R}^{T_w \times d_h} denote query, key, and value projections with head dimension dhd_h, and TwT_w is the sequence length per attention window.

    To limit computation on large sequence lengths, self-attention is computed within local windows: Stage 3 uses a window size of 14×1414 \times 14 tokens, and Stage 4 uses a window size of 7×77 \times 7 tokens.

  5. Knowl 5 — Ablation of Hybrid Integration Patterns of SSM and Self-Attention

    data/table

    Ablation study evaluating different hybrid placement configurations of Self-Attention blocks (SS) and MambaVision token mixer blocks (MM) in stages 3 and 4 of MambaVision-T (31.8M parameters, trained for 300 epochs on ImageNet-1K under iso-parameter constraints).

    Model Pattern Params (M) Top-1 (%)
    Random - 31.8 81.3
    First N/2N/2 layers SSSSMMMM 31.8 81.5
    Mixed layers-1 SMSMSMSM 31.8 81.4
    Mixed layers-2 MSMSMSMS 31.8 81.6
    Last N/4N/4 layers MMMMMMSS 31.8 81.9
    Last N/2N/2 layers MMMMSSSS 31.8 82.3

    Positioning self-attention blocks exclusively in the final N2\frac{N}{2} layers of stages 3 and 4 yields the highest accuracy (82.3% Top-1), outperforming interleaved patterns (81.4%−81.6%81.4\% - 81.6\%), front-loaded attention (81.5%81.5\%), and arbitrary random placement (81.3%81.3\%). This demonstrates that capturing global context via self-attention is most effective after hierarchical local feature extraction has been performed by earlier SSM and convolutional layers.

  6. Knowl 6 — Ablation on MambaVision Token Mixer Components

    data/table

    Systematic ablation analyzing each structural component of the MambaVision token mixer using the MambaVision-T architecture on ImageNet-1K classification, MS COCO object detection and instance segmentation (using Mask R-CNN with a 1×1\times schedule), and ADE20K semantic segmentation (using UPerNet).

    Configuration ImageNet Top-1 (%) COCO APbox\text{AP}^{\text{box}} COCO APmask\text{AP}^{\text{mask}} ADE20k mIoU (%)
    causal conv1 - w/o conv2 80.5 44.8 40.4 44.2
    conv1 - w/o conv2 80.9 45.0 40.8 44.7
    conv1 - conv2 - w/o concat 81.3 45.3 41.0 45.7
    conv1 - conv2 - concat 82.3 46.4 41.8 46.0

    conv1 denotes the 1D convolution in the SSM branch, conv2 denotes the 1D convolution in the parallel non-SSM branch, w/o conv2 indicates omitting the parallel branch (original Mamba design), and w/o concat uses Mamba's standard multiplicative gating mechanism instead of channel concatenation. Replacing causal convolution with regular convolution improves accuracy by +0.4% Top-1. Adding the symmetric non-SSM branch with gating adds another +0.4% Top-1. Replacing gating with concatenation provides a further +1.0% Top-1 gain (+1.8% overall), accompanied by +1.6 COCO APbox\text{AP}^{\text{box}}, +1.4 COCO APmask\text{AP}^{\text{mask}}, and +1.8 ADE20K mIoU.

  7. Knowl 7 — ImageNet-1K Classification Benchmarks across Backbone Families

    data/table

    Comparison of MambaVision model variants against representative convolutional, Vision Transformer, Conv-Transformer hybrid, and Mamba-based backbones on ImageNet-1K. Throughput is measured on a single NVIDIA A100 GPU with a batch size of 128.

    Model Image Size (Px) #Params (M) FLOPs (G) Throughput (Img/Sec) Top-1 (%)
    Conv-Based
    ConvNeXt-T 224 28.6 4.5 3196 82.0
    ConvNeXt-S 224 50.2 8.7 2008 83.1
    ConvNeXt-B 224 88.6 15.4 1485 83.8
    RegNetY-040 288 20.6 6.6 3227 83.0
    ResNetV2-101 224 44.5 7.8 4019 82.0
    EfficientNetV2-S 384 21.5 8.0 1735 83.9
    Transformer-Based
    Swin-T 224 28.3 4.4 2758 81.3
    Swin-S 224 49.6 8.5 1720 83.2
    SwinV2-T 256 28.3 4.4 1674 81.8
    SwinV2-S 256 49.7 8.5 1043 83.8
    SwinV2-B 256 87.9 15.1 535 84.6
    DeiT-B 224 86.6 16.9 2035 82.0
    DeiT3-L 224 304.4 59.7 535 84.8
    Conv-Transformer
    NextViT-S 224 31.7 5.8 3834 82.5
    NextViT-B 224 44.8 8.3 2926 83.2
    NextViT-L 224 57.8 10.8 2360 83.6
    MaxViT-B 224 120.0 23.4 507 84.9
    FasterViT-1 224 53.4 5.3 4188 83.2
    FasterViT-2 224 75.9 8.7 3161 84.2
    Mamba-Based
    Vim-T 224 7.0 - 3957 76.1
    Vim-S 224 26.0 - 1974 80.5
    EfficientVMamba-T 224 6.0 0.8 2904 76.5
    EfficientVMamba-S 224 11.0 1.3 1610 78.7
    EfficientVMamba-B 224 33.0 4.0 1482 81.8
    VMamba-T 224 30.0 4.9 1282 82.6
    VMamba-S 224 50.0 8.7 843 83.6
    VMamba-B 224 89.0 15.4 645 83.9
    MambaVision
    MambaVision-T 224 31.8 4.4 6298 82.3
    MambaVision-T2 224 35.1 5.1 5990 82.7
    MambaVision-S 224 50.1 7.5 4700 83.3
    MambaVision-B 224 97.7 15.0 3670 84.2
    MambaVision-L 224 227.9 34.9 2190 85.0
    MambaVision-L2 224 241.5 37.5 1021 85.3

    MambaVision achieves an improved accuracy-throughput Pareto frontier across all parameter tiers. In particular, MambaVision-B reaches 84.2% Top-1 at 3670 img/sec, outperforming VMamba-B (83.9% at 645 img/sec), Swin-S (83.2% at 1720 img/sec), and ConvNeXt-B (83.8% at 1485 img/sec).

  8. Knowl 8 — MS COCO Object Detection and Instance Segmentation Performance

    data/table

    Object detection and instance segmentation evaluation on the MS COCO dataset using Cascade Mask R-CNN with a 3×3\times schedule and an input crop resolution of 1280×8001280 \times 800.

    Backbone Params (M) FLOPs (G) APbox\text{AP}^{\text{box}} AP50box\text{AP}^{\text{box}}_{50} AP75box\text{AP}^{\text{box}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} AP75mask\text{AP}^{\text{mask}}_{75}
    DeiT-Small/16 80 889 48.0 67.2 51.7 41.4 64.2 44.3
    ResNet-50 82 739 46.3 64.3 50.5 40.1 61.7 43.4
    Swin-T 86 745 50.4 69.2 54.7 43.7 66.6 47.3
    ConvNeXt-T 86 741 50.4 69.1 54.8 43.7 66.5 47.3
    MambaVision-T 86 740 51.1 70.0 55.6 44.3 67.3 47.9
    X101-32 101 819 48.1 66.5 52.4 41.6 63.9 45.2
    Swin-S 107 838 51.9 70.7 56.3 45.0 68.2 48.8
    ConvNeXt-S 108 827 51.9 70.8 56.5 45.0 68.4 49.1
    MambaVision-S 108 828 52.3 71.1 56.7 45.2 68.5 48.9
    X101-64 140 972 48.3 66.4 52.3 41.7 64.0 45.1
    Swin-B 145 982 51.9 70.5 56.4 45.0 68.1 48.9
    ConvNeXt-B 146 964 52.7 71.3 57.2 45.6 68.9 49.5
    MambaVision-B 145 964 52.8 71.3 57.2 45.7 68.7 49.4

    MambaVision consistently outperforms comparably sized CNN (ResNet, ResNeXt, ConvNeXt) and Transformer (Swin, DeiT) backbones across Tiny, Small, and Base tiers in bounding box AP (APbox\text{AP}^{\text{box}}) and segmentation mask AP (APmask\text{AP}^{\text{mask}}).

  9. Knowl 9 — ADE20K Semantic Segmentation Performance

    data/table

    Semantic segmentation results on the ADE20K dataset using the UPerNet framework at an evaluation crop resolution of 512×512512 \times 512.

    Backbone Param (M) FLOPs (G) mIoU (%)
    DeiT-Small/16 52 1099 44.0
    Swin-T 60 945 44.5
    ResNet-101 86 1029 44.9
    Focal-T 62 998 45.8
    MambaVision-T 55 945 46.0
    Swin-S 81 1038 47.6
    Twins-SVT-B 89 - 47.7
    Focal-S 85 1130 48.0
    MambaVision-S 84 1135 48.2
    Swin-B 121 1188 48.1
    Twins-SVT-L 133 - 48.8
    Focal-B 126 1354 49.0
    MambaVision-B 126 1342 49.1

    MambaVision backbones outperform Swin Transformer baselines across all tiers (+1.5 mIoU for Tiny, +0.6 mIoU for Small, +1.0 mIoU for Base) as well as Focal Transformers and Twins-SVT backbones of comparable parameter and computational scale.

  10. Knowl 10 — ImageNet-21K Pre-training and High-Resolution Scaling

    empirical result

    Pre-training MambaVision on the ImageNet-21K dataset demonstrates effective scaling across model parameters and image resolutions:

    • MambaVision-B (97.7M params): Fine-tuning on ImageNet-1K at 224×224224 \times 224 resolution increases Top-1 accuracy from 84.2% (trained from scratch on ImageNet-1K) to 84.9%.
    • MambaVision-L (227.9M params): Pre-training on ImageNet-21K and fine-tuning on ImageNet-1K raises Top-1 accuracy from 85.0% to 86.1% at 224×224224 \times 224 resolution.
    • MambaVision-L3 (739.6M params): A scaled-up model achieves 87.3% Top-1 accuracy at 256×256256 \times 256 input resolution and 88.1% Top-1 accuracy when evaluated at 512×512512 \times 512 input resolution.

    These results establish that hybrid Mamba-Transformer vision architectures can successfully scale up to large dataset pre-training (~14M images), high parameter counts (~740M), and high input resolutions (512×512512 \times 512).

Coverage note — Qualitative attention map visualizations (Figure 5) demonstrating semantic object boundary localization were omitted as they provide illustrative rather than standalone quantitative or structural contributions.

References

  1. 1.Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE/CVF international conference on computer vision, pages 357–366, 2021. 2, 6
  2. 2.Zhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu, Longhui Wei, and Qi Tian. Visformer: The vision-friendly transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 589–598, 2021. 6
  3. 3.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems, 34, 2021. 2, 6, 7
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5, 6
  5. 5.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020. 2
  6. 6.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural networks, 107: 3–11, 2018. 4
  7. 7.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 1, 4, 5
  8. 8.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34:572–585, 2021. 3
  9. 9.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. Advances in Neural Information Processing Systems, 34:15908–15919, 2021. 6
  10. 10.Ali Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao, Jose M Alvarez, Jan Kautz, and Pavlo Molchanov. Fastervit: Fast vision transformers with hierarchical attention. arXiv preprint arXiv:2306.06189, 2023. 2, 6
  11. 11.Ali Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz, and Pavlo Molchanov. Global context vision transformers. In International Conference on Machine Learning, pages 12633–12646. PMLR, 2023. 5
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 2, 7
  13. 13.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 5, 6, 7, 8
  14. 14.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016. 3
  15. 15.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015. 3
  16. 16.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2
  17. 17.Jiashi Li, Xin Xia, Wei Li, Huixia Li, Xing Wang, Xuefeng Xiao, Rui Wang, Min Zheng, and Xin Pan. Next-vit: Next generation vision transformer for efficient deployment in realistic industrial scenarios. arXiv preprint arXiv:2207.05501, 2022. 2, 6
  18. 18.Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. Efficientformer: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems, 35:12934–12949, 2022. 2, 6
  19. 19.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 5, 7
  20. 20.Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, and Yunfan Liu. Vmamba: Visual state space model. arXiv preprint arXiv:2401.10166, 2024. 3, 6
  21. 21.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 2, 5, 6, 7
  22. 22.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022. 5, 6
  23. 23.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022. 2, 6, 7
  24. 24.Badri N Patro and Vijay S Agneeswaran. Simba: Simplified mamba-based architecture for vision and multivariate time series. arXiv preprint arXiv:2403.15360, 2024. 3, 6
  25. 25.Xiaohuan Pei, Tao Huang, and Chang Xu. Efficientvmamba: Atrous selective scan for light weight visual mamba. arXiv preprint arXiv:2403.09977, 2024. 1, 3, 6
  26. 26.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10428–10436, 2020. 2, 6
  27. 27.Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning, pages 10096–10106. PMLR, 2021. 2, 6
  28. 28.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021. 2, 6, 7
  29. 29.Hugo Touvron, Matthieu Cord, and Hervé Jégou. Deit iii: Revenge of the vit. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIV, pages 516–533. Springer, 2022. 6
  30. 30.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIV, pages 459–479. Springer, 2022. 6
  31. 31.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 1
  32. 32.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021. 2
  33. 33.Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021. 6
  34. 34.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), pages 418–434, 2018. 5, 6, 7
  35. 35.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1492–1500, 2017. 7
  36. 36.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9981–9990, 2021. 2, 6
  37. 37.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao. Focal attention for long-range interactions in vision transformers. Advances in Neural Information Processing Systems, 34, 2021. 5, 7
  38. 38.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10819–10829, 2022. 6
  39. 39.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017. 5, 6
  40. 40.Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. Vision mamba: Efficient visual representation learning with bidirectional state space model. In Proceedings of the 41st International Conference on Machine Learning, pages 62429–62442. PMLR, 2024. 1, 2, 6

Citation

MLA
Hatamizadeh, A., and J. Kautz. “MambaVision: A Hybrid Mamba-Transformer Vision Backbone”. arXiv, 2024, http://arxiv.org/abs/2407.08083v2.
APA
Hatamizadeh, A., & Kautz, J. (2024). MambaVision: A Hybrid Mamba-Transformer Vision Backbone. arXiv. http://arxiv.org/abs/2407.08083v2
Chicago
Hatamizadeh, A., and J. Kautz. 2024. “MambaVision: A Hybrid Mamba-Transformer Vision Backbone”. arXiv. http://arxiv.org/abs/2407.08083v2.
Harvard
Hatamizadeh, A. and Kautz, J. (2024) “MambaVision: A Hybrid Mamba-Transformer Vision Backbone”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.08083v2.
Vancouver
1. Hatamizadeh A, Kautz J (2024) MambaVision: A Hybrid Mamba-Transformer Vision Backbone. arXiv

BibTeX

@article{hatamizadeh2024mambavision,
  title = {MambaVision: A Hybrid Mamba-Transformer Vision Backbone},
  author = {Hatamizadeh, Ali and Kautz, Jan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.08083v2},
  eprint = {2407.08083}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE