AI Architecture & Systems

System Design for MoE Serving

5.23.1Expert Parallel Serving and Communication Paths#

5.23.2Combining Attention-DP with Expert Parallelism#

5.23.3Expert Parallel Load Balancing#

5.23.4Hot Expert Replication and Dynamic Placement#

5.23.5Expert Offload, Prefetch, and CPU–GPU Cooperation#

5.23.6Shared Experts, Routed Experts, and Cross-Device Execution#