AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Compiler
/
§3.12
Distributed Compilation
3.12.1
SPMD, Sharding Annotation, and Device Mesh
#
3.12.2
Automatic Parallelization Strategies and Graph Partitioning
#
3.12.3
Collective Insertion and Resharding
#
3.12.4
Compile-Time Scheduling of Communication–Computation Overlap
#
3.12.5
Pipeline Partitioning, Cross-Device Layout, and Memory Constraints
#
3.12.6
OpenXLA, GSPMD, and Shardy
#
← Previous
3.11 Cost Models, Autotuning, and Auto-Scheduling
→ Next
3.13 The PyTorch Compilation Stack