AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
System
/
§5.23
System Design for MoE Serving
5.23.1
Expert Parallel Serving and Communication Paths
#
5.23.2
Combining Attention-DP with Expert Parallelism
#
5.23.3
Expert Parallel Load Balancing
#
5.23.4
Hot Expert Replication and Dynamic Placement
#
5.23.5
Expert Offload, Prefetch, and CPU–GPU Cooperation
#
5.23.6
Shared Experts, Routed Experts, and Cross-Device Execution
#
← Previous
5.22 Routing and Distributed Serving Scheduling
→ Next
5.24 Deploying and Tuning Quantized Models