Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Mamba projection layers

Mamba projection layers are linear transformation components within Mamba selective state-space model architectures that map token representations between the model hidden dimension and expanded internal dimensions. Within a standard Mamba block, these layers include input projections that expand incoming representations into parallel paths for state-space sequence processing and multiplicative gating, internal projections that generate dynamic, input-dependent state-space matrices and step-size parameters, and a final output projection that maps the combined features back to the base model dimension. By governing these dimensional transitions and dynamic parameter projections, they facilitate selective information filtering and represent a substantial share of the trainable parameters and computational workload across state-space language models.

1 item

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Zheng Zhan, Liliang Ren, Shuohang Wang, Liyuan Liu, Yang Liu, Yeyun Gong, Yanzhi Wang, Yelong Shen

OrganizationsMicrosoftNortheastern University

Why you should read this

Proposes Routing Mamba, an architecture that scales state space models using sparse mixture-of-experts projections to match the language modeling performance of dense baselines while requiring 2.3 times fewer active parameters and cutting compute costs by 23 percent.

Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3x more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance.

Added

2026-10-01