Sparse Mixture of Experts

Sparse Mixture of Experts (SMoE) is a neural network architecture where the model is composed of multiple “expert” sub-networks. For each input token, a gating mechanism selects a small subset of experts to process the data, while the remaining experts remain inactive. This allows for massive model capacity with significantly reduced computational cost during inference, as only a fraction of parameters are activated per token.

Core Mechanism

  • Expert Routing: A router network determines which experts are most relevant for a given input.
  • Sparsity: Only the top- experts (typically or ) are activated per token, keeping latency low despite the large total parameter count.
  • Load Balancing: Techniques are employed to ensure experts are utilized evenly to prevent bottlenecks.

Hardware Efficiency & Inference

SMoE architectures are particularly suited for deployment on consumer-grade hardware through advanced runtime optimizations that leverage MoE sparsity.

References

Source Notes