AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.12
MoE Kernels
2.12.1
Router, Top-k Gating, and Token Assignment
#
2.12.2
Token Permutation and Unpermutation
#
2.12.3
Grouped GEMM and Expert Padding
#
2.12.4
Dropless MoE and Block-Sparse GEMM
#
2.12.5
Dispatch/Combine and Communication Fusion
#
2.12.6
Shared Experts, Routed Experts, and Compute Overlap
#
← Previous
2.11 Sparse and Irregular Kernels
→ Next
2.13 Low-Precision and Quantized Kernels