Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Zheng ZhanLiliang RenShuohang WangLiyuan LiuYang LiuYeyun GongYanzhi WangYelong Shen

article2025NeurIPS4 citations

Proposes Routing Mamba, an architecture that scales state space models using sparse mixture-of-experts projections to match the language modeling performance of dense baselines while requiring 2.3 times fewer active parameters and cutting compute costs by 23 percent.

Listen

Modern natural language processing relies heavily on scalable architectures, but traditional models face severe computational and memory bottlenecks when handling long sequences. While State Space Models such as Mamba offer efficient sequence processing with constant-time inference, efficiently scaling their capacity has remained an unresolved challenge. In traditional architectures, Mixture of Experts designs scale model capacity by sparsely activating specialized sub-networks for each input token. However, naive attempts to apply Mixture of Experts to Mamba layers fail, leading to performance degradation and increased latency due to uncoordinated routing across interdependent internal layers.

The article demonstrates an effective method to scale State Space Models using sparse mixtures of linear projection experts. Specifically, it introduces Routing Mamba, a framework designed to expand model capacity and improve sequence modeling performance while maintaining computational efficiency.

To achieve this, the authors evaluated language models ranging from 115 million to 10 billion total parameters using the SlimPajama dataset across sequence lengths up to 16,000 tokens. The technical approach targets the computationally heavy projection layers within Mamba—namely the Convolution, Gate, and Output projections—converting them into sparsely activated expert pools where only one out of eight experts is activated per token. Crucially, the system uses a shared routing decision across these projection layers, ensuring that all projections for a given token act in harmony. Smaller, specialized sub-components share parameters across experts, and the architecture functions stably without requiring artificial load-balancing loss penalties.

The key findings show significant improvements in efficiency and model quality. First, Routing Mamba achieves language modeling performance equivalent to dense Mamba models that require up to 2.3 times more active parameters. Second, when applied to hybrid architectures that blend Mamba with attention mechanisms, the framework reduces floating-point operations by 23% compared to dense layer expansion while maintaining equivalent accuracy. Third, the model demonstrates robust context-length generalization, consistently maintaining lower perplexity scores than baseline models across extended sequence lengths. Finally, the framework successfully scaled other state space variants, including Mamba-2 and Gated DeltaNet, while achieving about 80% relative training throughput compared to dense baselines despite having more than twice the total parameter count.

These results demonstrate that sparse expert routing can be extended beyond conventional feed-forward layers to state space projections without incurring significant computational overhead. For organizations deploying large language models, this approach allows for higher accuracy and larger functional model capacity at a fraction of the inference compute cost, translating to reduced hardware expenditures and lower operational latency.

Organizations developing sequence modeling systems should consider adopting shared-routing expert architectures when scaling state space or hybrid models. In practice, practitioners should selectively expertize major projection layers while keeping smaller auxiliary components shared. Further exploration is recommended to test optimal configurations across larger multi-billion-parameter scales and emerging state space variants before wide-scale deployment. Although confidence in these empirical results is high across the tested benchmarks, readers should note that boundary conditions remain unmapped for non-language modalities and diverse linear attention designs.

arXiv: 2506.18145

No sufficiently relevant recommendations were found.

Cover for Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Abstract

Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3x more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Mamba Layer
  • 3.2 Mixture of Experts
  • 4 Design of Routing Mamba
  • 4.1 The Challenge of Naive MoE Integration in Mamba
  • 4.2 RoM Methodology
  • 4.3 Details of Design Choices
  • 5 Experiments
  • 5.1 Implementation Details
  • 5.2 Scaling on Mamba
  • 5.3 Comparison and Ablation Study
  • 5.4 Empirical Insights on Applying RoM-style MoE Across Architectures
  • 6 Conclusion
  • 7 Acknowledgments
  • References
  • A Appendix
  • A.1 Detailed Comparisons Results for Potential RoM Strategies
  • A.2 Architectures Overview
  • A.3 Load Balance and More Router Details
  • A.4 More Experimental Results
  • A.5 More experimental Results of Hybrid RoM + FFN-MoE
  • A.6 Training Throughput Results of RoM

Knowls

  1. Knowl 1 — RoM uses one sparse routing decision for Mamba projection experts

    model/method

    Routing Mamba (RoM) turns Mamba’s input, gate, and output projections into expertized linear layers and uses one router to select the same expert pathway for all three. For a token representation xt∈RDmx_t\in\mathbb{R}^{D_m}, let NN be the number of experts, KK the number selected, and Wr∈RDm×NW_r\in\mathbb{R}^{D_m\times N} the router weights. The router forms probabilities pt=Softmax⁡(xtWr)p_t=\operatorname{Softmax}(x_tW_r) and selects the index set StS_t of the KK largest probabilities; inactive experts receive zero routing weight, with rt,i=pt,i1[i∈St]r_{t,i}=p_{t,i}\mathbf{1}[i\in S_t].

    Let Win,i,Wg,i∈RDm×DeW_{in,i},W_{g,i}\in\mathbb{R}^{D_m\times D_e} and Wout,i∈RDe×DmW_{out,i}\in\mathbb{R}^{D_e\times D_m} be the input, gate, and output projection matrices for expert ii. RoM computes the input and gate projections using the shared selected set:

    Ht=∑i=1N1[i∈St]xtWin,i,Gt=SiLU⁡ ⁣(∑i=1N1[i∈St]xtWg,i).H_t=\sum_{i=1}^{N}\mathbf{1}[i\in S_t]x_tW_{in,i},\qquad G_t=\operatorname{SiLU}\!\left(\sum_{i=1}^{N}\mathbf{1}[i\in S_t]x_tW_{g,i}\right).

    The short convolution and selective state-space computation applied to HtH_t produce Yt∈RDeY_t\in\mathbb{R}^{D_e}. The block output is

    Ot=∑i=1Nrt,i(Yt⊙Gt)Wout,i,O_t=\sum_{i=1}^{N}r_{t,i}(Y_t\odot G_t)W_{out,i},

    where ⊙\odot denotes elementwise multiplication. Thus the same token-level selection coordinates the three functionally linked projections, rather than routing each projection independently. In the main RoM experiments, one of eight experts is activated per token at each layer. The authors’ proposed rationale is that a unified expert pathway encourages coherent token specialization and reduces the separate routers’ learning burden.

  2. Knowl 2 — RoM lowers perplexity at fixed active Mamba model sizes

    empirical result

    On SlimPajama language modeling, RoM improves validation perplexity (PPL; lower is better) over dense Mamba at each tested active parameter scale. The following results are for models trained on 20B tokens with a 4K training sequence length; validation PPL is shown at four evaluation context lengths. RoM activates one of eight experts per token, so its active parameter count matches the dense model even though its total parameter count is larger.

    ArchitectureActive parametersTotal parametersPPL at 4,096PPL at 8,192PPL at 12,288PPL at 16,384
    Mamba115M115M14.3113.8913.7113.70
    RoM115M710M12.3311.9711.8211.79
    Mamba353M353M11.1810.8310.6810.65
    RoM353M2.5B10.009.699.559.52
    Mamba765M765M9.849.529.379.34
    RoM765M5.5B8.958.688.558.52
    Mamba1.3B1.3B9.078.768.638.60
    RoM1.3B10B8.528.278.158.12

    These comparisons show improved PPL at matched active size across all four scales, with RoM’s larger total parameter count sparsely activated. The authors also report that, in their scaling curves, dense Mamba required up to 2.3 times the active parameters to reach comparable PPL to RoM.

  3. Knowl 3 — RoM matches a wider Samba model with fewer FLOPs

    empirical result

    In language-model validation on SlimPajama, RoM improves Samba’s PPL while retaining its active parameter count. The comparison below reports active and total parameters, forward-pass FLOPs at sequence length 4K, and validation PPL at context lengths 4K, 8K, and 16K. The page 8 results grid makes the key comparison visible: Samba with RoM (Conv, Gate, Out) has nearly the same PPL as dense Samba with expand=4, while using 23% fewer FLOPs.

    ArchitectureActive parametersTotal parametersFLOPsPPL at 4KPPL at 8KPPL at 16K
    Samba (expand=2)421M421M4.74T10.059.649.56
    Samba + MoA421M1.1B4.74T9.949.549.46
    Samba + SwitchHead421M1.1B4.74T9.849.459.37
    Samba + independent MoE-Mamba (Conv, Gate, Out)421M1.0B4.74T10.269.859.77
    Samba + RoM (Conv, Gate, Out)421M1.0B4.74T9.779.389.31
    Samba (expand=4)511M511M6.13T9.789.399.31
    Samba + RoM (Gate, Out)511M1.3B6.13T9.379.008.93
    Samba + RoM (Conv, Gate, Out)511M1.7B6.13T9.288.928.85
    Samba + RoM (Conv, Gate, dt, x, Out)511M1.7B6.13T9.308.938.87

    In particular, the 421M-active RoM model uses 4.74T FLOPs versus 6.13T for dense Samba (expand=4), a 23% reduction, and gives PPL 9.77 versus 9.78 at 4K and 9.31 versus 9.31 at 16K. RoM also outperforms the listed attention-projection MoE baselines at all three validation contexts.

  4. Knowl 4 — Independent MoE routing degrades Mamba projection performance

    empirical result

    A comparison on Samba with 421M active parameters tested conventional independent MoE routing on Mamba’s Conv, Gate, and Output projections, with one of eight experts active per token. All models were trained on 20B SlimPajama tokens using 8×A100 GPUs. Validation PPL is reported at context lengths 4K, 8K, and 16K.

    ConfigurationTotal parametersPPL at 4KPPL at 8KPPL at 16K
    Dense Samba baseline421M10.059.649.56
    Independent MoE on Conv620M10.069.669.58
    Independent MoE on Gate620M10.109.699.60
    Independent MoE on Output620M10.149.739.65
    Independent MoE on Conv and Gate818M10.129.709.62
    Independent MoE on Conv and Output818M10.179.759.68
    Independent MoE on Gate and Output818M10.189.779.69
    Independent MoE on Conv, Gate, and Output1.0B10.269.859.77
    RoM on Conv, Gate, and Output1.0B9.779.389.31

    Every independent-routing variant is worse than the dense baseline at each validation context, and the degradation grows when more projections are expertized. At the same 1.0B total parameter count, shared-routing RoM instead improves over the dense baseline. The plotted comparison on page 5 likewise highlights the 4K PPL difference between the dense baseline (10.05) and RoM (9.77).

  5. Knowl 5 — RoM retains its advantage across training and evaluation context lengths

    empirical result

    The paper evaluates whether RoM’s PPL improvement persists when training and validation sequence lengths differ. The following 1.3B-active-parameter results use 20B training tokens and report validation PPL at 4K, 8K, 12K, and 16K contexts. RoM has 10B total parameters and activates one of eight experts per token; the dense Mamba baseline has 1.3B total parameters.

    Training sequence lengthArchitecturePPL at 4KPPL at 8KPPL at 12KPPL at 16K
    4KMamba9.078.768.638.60
    4KRoM8.528.278.158.12
    8KMamba9.168.868.718.67
    8KRoM8.528.258.128.08
    16KMamba9.198.878.728.68
    16KRoM8.658.388.238.19

    RoM has lower PPL than dense Mamba at every listed validation context for all three training sequence lengths. This supports the paper’s reported consistency across training contexts and its ability to evaluate at longer contexts than those used in training.

  6. Knowl 6 — Shared-routing RoM also improves other linear recurrent models

    empirical result

    The paper applies RoM-style shared routing beyond the original Mamba architecture and reports SlimPajama validation PPL at contexts of 4K, 8K, 12K, and 16K. The table gives active and total parameters for each architecture.

    ArchitectureActive parametersTotal parametersPPL at 4KPPL at 8KPPL at 12KPPL at 16K
    Mamba353M353M11.1810.8310.6810.65
    Mamba + RoM353M2.5B10.009.699.559.52
    Mamba2 + RoM352M2.5B9.829.499.349.30
    Gated DeltaNet + RoM343M2.5B9.819.469.299.26

    RoM lowers validation PPL relative to the reported dense Mamba baseline, and the RoM-configured Mamba2 and Gated DeltaNet models also obtain lower PPL than the RoM-configured original Mamba at the listed contexts. The authors report that, for the more unified Mamba2 and Gated DeltaNet blocks, expertizing the major projection pathways comprehensively works well.

  7. Knowl 7 — RoM combines with feed-forward MoE at comparable language-model PPL

    empirical result

    The paper tests adding RoM to Samba models that also use feed-forward MoE (FFN-MoE). The results below report active and total parameters and SlimPajama validation PPL at 4K, 8K, and 16K contexts. The two training settings differ in token count, microbatch size, and GPU count, as shown.

    Training settingArchitectureActive parametersTotal parametersPPL at 4KPPL at 8KPPL at 16K
    20B tokens; microbatch 4; 8×A100Samba + FFN-MoE (16 top-1)511M2.8B8.808.468.40
    20B tokens; microbatch 4; 8×A100Samba + RoM + FFN-MoE (8 top-1)511M2.9B8.838.498.43
    17B tokens; microbatch 1; 16×A100Samba + FFN-MoE (32 top-1)511M5.7B8.888.608.45
    17B tokens; microbatch 1; 16×A100Samba + RoM + FFN-MoE (16 top-1)511M5.6B8.988.708.54

    In the 20B-token setting, the hybrid model is within 0.03 PPL of the FFN-MoE baseline at each context while adding 0.1B total parameters. In the 17B-token setting, it has 0.1B fewer total parameters and is within 0.10 PPL at each context of the larger FFN-MoE model.

  8. Knowl 8 — Hybrid RoM and FFN-MoE preserves downstream task accuracy

    empirical result

    On Samba with 511M active parameters, the paper compares FFN-MoE with hybrid RoM + FFN-MoE on common-sense reasoning tasks. The table reports LAMBADA perplexity and accuracy, accuracy on HellaSwag, PIQA, ARC-Easy, and WinoGrande, and the mean accuracy across the five accuracy metrics. Expert counts are indicated by top-1 routing; total parameters vary with the expert configuration.

    MethodTotal parametersLAMBADA PPLLAMBADA accuracy (%)HellaSwag (%)PIQA (%)ARC-Easy (%)WinoGrande (%)Mean accuracy (%)
    FFN-MoE (16 top-1)2.8B24.838.442.168.150.152.350.2
    RoM + FFN-MoE (8 top-1)2.9B24.839.541.668.148.752.850.1
    FFN-MoE (32 top-1)5.7B26.938.040.767.450.351.149.5
    RoM + FFN-MoE (16 top-1)5.6B28.536.441.068.048.452.649.2

    The 2.9B hybrid configuration has a mean accuracy of 50.1%, close to 50.2% for the 2.8B FFN-MoE baseline. At the larger scale, the 5.6B hybrid model reaches 49.2% mean accuracy versus 49.5% for the 5.7B FFN-MoE baseline; task-level results are mixed rather than uniformly favoring either configuration.

  9. Knowl 9 — Sharing small Mamba components is preferable to expertizing them

    empirical result

    For original Mamba, RoM expertizes the larger Conv, Gate, and Output projections but shares the smaller x projection, dt projection, and 1D convolution across experts. In a Samba expand=4 comparison with 511M active parameters, both RoM variants below have 1.7B total parameters and 6.13T forward-pass FLOPs at sequence length 4K.

    Expertized projectionsPPL at 4KPPL at 8KPPL at 16K
    Conv, Gate, Output9.288.928.85
    Conv, Gate, dt, x, Output9.308.938.87

    Adding the smaller dt and x components to the expertized set gives slightly higher PPL at all three validation contexts. This ablation supports the design choice to share small specialized components rather than replicate them across experts.

  10. Knowl 10 — RoM’s training throughput exceeds a wider dense Samba baseline

    empirical result

    Training throughput was measured on 8×A100 GPUs for Samba models trained on 20B tokens. The reported rate is in units of 10510^5 tokens per second.

    ArchitectureActive parametersTotal parametersTraining speed (×105\times10^5 tokens/s)
    Samba (expand=2)421M421M4.42
    Samba + RoM (Conv, Gate, Out)421M1.0B3.57
    Samba (expand=4)511M511M3.34

    The RoM model processes about 80% as many tokens per second as the expand=2 dense model despite having over twice its total parameters. Its measured throughput is also higher than that of the wider expand=4 dense model, while the separate validation comparison reports similar PPL to expand=4 with fewer FLOPs.

  11. Knowl 11 — Explicit load-balancing loss does not materially improve RoM

    empirical result

    The paper evaluates whether an auxiliary load-balancing loss is needed to distribute tokens across experts. For Samba expand=4 with RoM on Conv, Gate, and Output, both configurations have 511M active and 1.7B total parameters. Validation PPL without versus with the load-balancing loss was 9.28 versus 9.30 at 4K, 8.92 versus 8.94 at 8K, and 8.85 versus 8.87 at 16K. For the configuration that also expertizes dt and x, the corresponding values were 9.30 versus 9.26, 8.93 versus 8.89, and 8.87 versus 8.83. The small differences provide no consistent evidence of an improvement from the auxiliary loss; the paper therefore uses RoM without it by default.

  12. Knowl 12 — Language-model experiments use a consistent SlimPajama training protocol

    experimental setup

    Unless otherwise specified, the paper trains its language models on SlimPajama for 20B tokens with a 4K sequence length and reports validation perplexity; some experiments vary training context length, token budget, microbatch size, or hardware and identify those changes alongside their results. The common optimizer is AdamW with β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, gradient clipping at 1.0, weight decay 0.1, a cosine learning-rate schedule, maximum learning rate 4×10−44\times10^{-4}, and warmup ratio 0.01. The default global batch contains 2 million tokens.

    Training uses PyTorch Fully Sharded Data Parallel with CPU offloading. The experiments avoid expert parallelism, so they do not require a capacity factor or token dropping; Megablocks grouped GEMM kernels are used to accelerate MoE computation. The usual RoM setting activates one of eight experts per token per layer.

  13. Knowl 13 — RoM’s optimal configuration and broader applicability remain uncertain

    limitation

    The paper identifies uncertainty about the best RoM configuration and about how broadly its approach transfers across the evolving range of state-space model variants. Its applicability to self-attention and linear-attention architectures is also unresolved. The reported experiments provide evidence for selected Mamba-style models and hybrid configurations, but do not establish a universally optimal expertization strategy for those broader architecture classes.

Coverage note — The paper’s router jitter and SparseMixer gradient-estimation details are not separate knowls because they are training aids rather than standalone contributions; the reported evaluation results, design ablations, and stated limitations are included.

References

  1. 1.Quentin Anthony, Yury Tokpanov, Paolo Glorioso, and Beren Millidge. Blackmamba: Mixture of experts for state-space models. arXiv preprint arXiv:2402.01771, 2024.
  2. 2.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020.
  3. 3.François Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017.
  4. 4.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  5. 5.R’obert Csord’as, Piotr Piekos, Kazuki Irie, and Jürgen Schmidhuber. Switchhead: Accelerating transformers with mixture-of-experts attention. ArXiv, abs/2312.07987, 2023.
  6. 6.Tri Dao and Albert Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality. arXiv preprint arXiv:2405.21060, 2024.
  7. 7.Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. Language modeling with gated convolutional networks. In International conference on machine learning, pages 933–941. PMLR, 2017.
  8. 8.Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre. Griffin: Mixing gated linear recurrences with local attention for efficient language models. arXiv preprint arXiv: 2402.19427, 2024.
  9. 9.DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv: 2412.19437, 2024.
  10. 10.Xin Dong, Y. Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Maheshwary, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, Yingyan Celine Lin, Jan Kautz, and Pavlo Molchanov. Hymba: A hybrid-head architecture for small language models. ArXiv, abs/2411.13676, 2024.
  11. 11.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pages 5547–5569. PMLR, 2022.
  12. 12.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural networks, 107:3–11, 2018.
  13. 13.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022.
  14. 14.Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts. Proceedings of Machine Learning and Systems, 5, 2023.
  15. 15.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
  16. 16.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
  17. 17.Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. On the parameterization and initialization of diagonal state space models. Advances in Neural Information Processing Systems, 35:35971–35983, 2022.
  18. 18.Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  19. 19.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34:572–585, 2021.
  20. 20.Ankit Gupta, Albert Gu, and Jonathan Berant. Diagonal state spaces are as effective as structured state spaces. Advances in Neural Information Processing Systems, 35:22982–22994, 2022.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  22. 22.Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. Adaptive mixtures of local experts. Neural computation, 3(1):79–87, 1991.
  23. 23.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  24. 24.Rudolph Emil Kalman. A new approach to linear filtering and prediction problems. 1960.
  25. 25.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  26. 26.Weixin Liang, Junhong Shen, Genghan Zhang, Ning Dong, Luke S. Zettlemoyer, and Lili Yu. Mixture-of-mamba: Enhancing multi-modal state-space models with modality-aware sparsity. 2025.
  27. 27.Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Haim Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avshalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham. Jamba: A hybrid transformer-mamba language model. ArXiv, abs/2403.19887, 2024.
  28. 28.Liyuan Liu, Chengyu Dong, Xiaodong Liu, Bin Yu, and Jianfeng Gao. Bridging discrete and backpropagation: Straight-through and beyond. In arXiv:2304.08612 [cs], 2023.
  29. 29.Liyuan Liu, Jianfeng Gao, and Weizhu Chen. Sparse backpropagation for moe training. arXiv preprint arXiv:2310.00811, 2023.
  30. 30.Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou. Megalodon: Efficient llm pretraining and inference with unlimited context length. Advances in Neural Information Processing Systems, 37:71831–71854, 2024.
  31. 31.Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. Long range language modeling via gated state spaces. arXiv preprint arXiv:2206.13947, 2022.
  32. 32.MiniMax. Minimax-01: Scaling foundation models with lightning attention. arXiv preprint arXiv: 2501.08313, 2025.
  33. 33.Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal. Leave no context behind: Efficient infinite context transformers with infini-attention. arXiv preprint arXiv:2404.07143, 101, 2024.
  34. 34.OpenAI. Gpt-4 technical report. PREPRINT, 2023.
  35. 35.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  36. 36.Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee, and Dimitris Papailiopoulos. Can mamba learn how to learn? a comparative study on in-context learning tasks. arXiv preprint arXiv:2402.04248, 2024.
  37. 37.Maciej Pi’oro, Kamil Ciebiera, Krystian Kr’ol, Jan Ludziejewski, and Sebastian Jaszczur. Moe-mamba: Efficient selective state space models with mixture of experts. ArXiv, abs/2401.04081, 2024.
  38. 38.Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel J Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. Hyena hierarchy: Towards larger convolutional language models. In International Conference on Machine Learning, pages 28043–28078. PMLR, 2023.
  39. 39.Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen. Samba: Simple hybrid state space models for efficient unlimited context language modeling. arXiv preprint arXiv:2406.07522, 2024.
  40. 40.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  41. 41.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  42. 42.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  43. 43.Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 2023.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  45. 45.Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, and Raffay Hamid. Selective structured state-spaces for long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6387–6397, 2023.
  46. 46.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  47. 47.Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. arXiv preprint arXiv:2412.06464, 2024.
  48. 48.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  49. 49.Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  50. 50.Xiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou, Wenge Rong, and Zhang Xiong. Mixture of attention heads: Selecting attention heads per token. In Conference on Empirical Methods in Natural Language Processing, 2022.

Citation

MLA
Zhan, Z., et al. “Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection”. arXiv, 2025, http://arxiv.org/abs/2506.18145v1.
APA
Zhan, Z., Ren, L., Wang, S., Liu, L., Liu, Y., Gong, Y., Wang, Y., & Shen, Y. (2025). Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection. arXiv. http://arxiv.org/abs/2506.18145v1
Chicago
Zhan, Z., L. Ren, S. Wang, et al. 2025. “Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection”. arXiv. http://arxiv.org/abs/2506.18145v1.
Harvard
Zhan, Z. et al. (2025) “Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.18145v1.
Vancouver
1. Zhan Z, Ren L, Wang S, Liu L, Liu Y, Gong Y, Wang Y, Shen Y (2025) Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection. arXiv

BibTeX

@article{zhan2025routing,
  title = {Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection},
  author = {Zhan, Zheng and Ren, Liliang and Wang, Shuohang and Liu, Liyuan and Liu, Yang and Gong, Yeyun and Wang, Yanzhi and Shen, Yelong},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.18145v1},
  eprint = {2506.18145}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/