AI Architecture and Systems Tutorial
A topic map from GEMM execution to datacenter-scale AI, centered on generality, specialization, and data movement.
This outline connects six layers: Microarchitecture → Kernel → Compiler → Architecture → System → Algorithms. Hardware chapters follow the supplied survey, Balancing Generality and Specialization: A Survey on AI Datacenter Hardware Architecture, and the AI Datacenter Accelerator Research Corpus. Kernel, compiler, and runtime comparisons use the corpus’s per-chip software mappings. The broader systems and algorithms topics are carried forward from the original outline, with Awesome-ML-SYS-Tutorial as a systems reading index.
Source map and scope
Survey. Yufeng Gu, Jiazhen Wang, and Reetuparna Das, September 2026, supplied 34-page draft. “Survey §…” and figure references refer to that attachment. Companion project page.
Corpus. Platform links are pinned to commit c49a55c6cdbb. Per-chip records distinguish confirmed, inferred, contested, nonpublic, historical, and announced information. Those qualifications remain relevant when developing the corresponding chapters.
Core and extensions. The survey’s core taxonomy is GPU, NPU, Spatial Dataflow, and Compute-in-Memory. Additional vendor cases, photonic/neuromorphic comparisons, and CXL/HBF topics are labeled extensions. Untagged foundations and the detailed systems/algorithms curriculum are retained topics, rather than findings attributed to the hardware survey.
Cross-layer boundaries. Numerical hardware support, quantized kernels, compiler transformations, deployment policy, and quantization algorithms are separate topics. Likewise, collective semantics, communication algorithms, and physical topology are separated. A physical server/rack/pod need not coincide with a scale-up domain.