AI Architecture & Systems

Mixture of Experts

6.10.1Dense FFN → Sparse MoE#

6.10.2Router, Top-k Routing, and Expert Selection#

6.10.3Token Choice, Expert Choice, and Capacity Constraints#

6.10.4Shared Experts, Fine-Grained Experts, and Expert Partitioning#

6.10.5Auxiliary Loss and Auxiliary-Loss-Free Balancing#

6.10.6Expert Specialization, Routing Collapse, and Training Stability#

6.10.7Activated Parameters, Total Parameter Count, and Effective Compute#