AI Architecture & Systems

Distributed Training Frameworks and Accelerator Backends

SourcesOriginal training outline, extended with TPU, Neuron and Ascend framework mappings.

5.16.1PyTorch Distributed: DDP, FSDP, and DTensor#

5.16.2Megatron-LM/Megatron-Core and DeepSpeed#

5.16.3TorchTitan and Native PyTorch Training#

5.16.4JAX/MaxText and TPU Training#

5.16.5NeMo and Training Workflow Integration#

5.16.6AWS NxD Training, Ascend Adaptation, and Device-Specific Training Paths#

5.16.7Distinguishing Unified Training Concepts from Per-Backend Support Coverage#