AI Architecture & Systems

Data Parallelism and Parameter Sharding

5.6.1Synchronous SGD and Distributed Data Parallel#

5.6.2Gradient Bucketing and Communication Overlap#

5.6.3ZeRO-1/2/3 and State Sharding#

5.6.4FSDP/FSDP2 and Parameter All-Gather#

5.6.5Gradient Accumulation and Effective Batch Size#

5.6.6Replicated, Sharded, and Hybrid Sharded Training#