Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
Zheng ZhanLiliang RenShuohang WangLiyuan LiuYang LiuYeyun GongYanzhi WangYelong Shen
Proposes Routing Mamba, an architecture that scales state space models using sparse mixture-of-experts projections to match the language modeling performance of dense baselines while requiring 2.3 times fewer active parameters and cutting compute costs by 23 percent.
Modern natural language processing relies heavily on scalable architectures, but traditional models face severe computational and memory bottlenecks when handling long sequences. While State Space Models such as Mamba offer efficient sequence processing with constant-time inference, efficiently scaling their capacity has remained an unresolved challenge. In traditional architectures, Mixture of Experts designs scale model capacity by sparsely activating specialized sub-networks for each input token. However, naive attempts to apply Mixture of Experts to Mamba layers fail, leading to performance degradation and increased latency due to uncoordinated routing across interdependent internal layers.
The article demonstrates an effective method to scale State Space Models using sparse mixtures of linear projection experts. Specifically, it introduces Routing Mamba, a framework designed to expand model capacity and improve sequence modeling performance while maintaining computational efficiency.
To achieve this, the authors evaluated language models ranging from 115 million to 10 billion total parameters using the SlimPajama dataset across sequence lengths up to 16,000 tokens. The technical approach targets the computationally heavy projection layers within Mamba—namely the Convolution, Gate, and Output projections—converting them into sparsely activated expert pools where only one out of eight experts is activated per token. Crucially, the system uses a shared routing decision across these projection layers, ensuring that all projections for a given token act in harmony. Smaller, specialized sub-components share parameters across experts, and the architecture functions stably without requiring artificial load-balancing loss penalties.
The key findings show significant improvements in efficiency and model quality. First, Routing Mamba achieves language modeling performance equivalent to dense Mamba models that require up to 2.3 times more active parameters. Second, when applied to hybrid architectures that blend Mamba with attention mechanisms, the framework reduces floating-point operations by 23% compared to dense layer expansion while maintaining equivalent accuracy. Third, the model demonstrates robust context-length generalization, consistently maintaining lower perplexity scores than baseline models across extended sequence lengths. Finally, the framework successfully scaled other state space variants, including Mamba-2 and Gated DeltaNet, while achieving about 80% relative training throughput compared to dense baselines despite having more than twice the total parameter count.
These results demonstrate that sparse expert routing can be extended beyond conventional feed-forward layers to state space projections without incurring significant computational overhead. For organizations deploying large language models, this approach allows for higher accuracy and larger functional model capacity at a fraction of the inference compute cost, translating to reduced hardware expenditures and lower operational latency.
Organizations developing sequence modeling systems should consider adopting shared-routing expert architectures when scaling state space or hybrid models. In practice, practitioners should selectively expertize major projection layers while keeping smaller auxiliary components shared. Further exploration is recommended to test optimal configurations across larger multi-billion-parameter scales and emerging state space variants before wide-scale deployment. Although confidence in these empirical results is high across the tested benchmarks, readers should note that boundary conditions remain unmapped for non-language modalities and diverse linear attention designs.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Read the original Mamba paper first to understand the selective state-space block and hardware-aware scan that Routing Mamba modifies.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). This foundational sparse-gating design explains the expert routing mechanism that Routing Mamba adapts for projection layers.
- Paper: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, William Fedus et al. (2022). Switch Transformers provides a concrete language-modeling precedent for sparse expert routing and scaling total capacity while limiting active computation.
- Paper: Unified Scaling Laws for Routed Language Models, Aidan Clark et al. (2022). Its routed-language-model scaling laws clarify how active size and expert count interact, framing the scaling claims RoM tests for SSMs.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). S4 supplies the structured state-space foundations behind later efficient SSM architectures, including the Mamba family RoM builds on.
No sufficiently relevant recommendations were found.
