Built independently by an author, for readers. Read the story and support ChapterPal

keyword

mixture-of-experts projection

A mixture-of-experts projection is a neural network transformation layer that replaces a standard dense linear projection with a collection of specialized sub-projection matrices, known as experts, whose activations are dynamically determined by a routing mechanism. In this architecture, an input-dependent gating function evaluates incoming token representations and directs them to a sparse subset of available projection experts rather than processing all inputs through a single shared weight matrix. This design expands the total parameter capacity and representational richness of the network while keeping active computational cost bounded during execution. By applying sparse conditional routing directly to linear projection operations, mixture-of-experts projections facilitate efficient parameter scaling in deep learning architectures, such as state space models and hybrid sequence models, without incurring the computational overhead of scaling dense projection layers.

1 item

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

Zheng Zhan, Liliang Ren, Shuohang Wang, Liyuan Liu, Yang Liu, Yeyun Gong, Yanzhi Wang, Yelong Shen

OrganizationsMicrosoftNortheastern University

Why you should read this

Proposes Routing Mamba, an architecture that scales state space models using sparse mixture-of-experts projections to match the language modeling performance of dense baselines while requiring 2.3 times fewer active parameters and cutting compute costs by 23 percent.

Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3x more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance.

Added

2026-10-01