AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
System
/
§5.6
Data Parallelism and Parameter Sharding
5.6.1
Synchronous SGD and Distributed Data Parallel
#
5.6.2
Gradient Bucketing and Communication Overlap
#
5.6.3
ZeRO-1/2/3 and State Sharding
#
5.6.4
FSDP/FSDP2 and Parameter All-Gather
#
5.6.5
Gradient Accumulation and Effective Batch Size
#
5.6.6
Replicated, Sharded, and Hybrid Sharded Training
#
← Previous
5.5 Distributed Runtime and Collective Communication
→ Next
5.7 Tensor Parallelism