keyword
mixture-of-experts projection
A mixture-of-experts projection is a neural network transformation layer that replaces a standard dense linear projection with a collection of specialized sub-projection matrices, known as experts, whose activations are dynamically determined by a routing mechanism. In this architecture, an input-dependent gating function evaluates incoming token representations and directs them to a sparse subset of available projection experts rather than processing all inputs through a single shared weight matrix. This design expands the total parameter capacity and representational richness of the network while keeping active computational cost bounded during execution. By applying sparse conditional routing directly to linear projection operations, mixture-of-experts projections facilitate efficient parameter scaling in deep learning architectures, such as state space models and hybrid sequence models, without incurring the computational overhead of scaling dense projection layers.
1 item

