AI Architecture & Systems

GPU: General-Purpose Parallel Execution and Specialized Matrix Computation

SourcesSurvey §3 and §3.1; additional GPU records are corpus extensions, not additional survey case studies.

4.4.1NVIDIA GPU: SM, Tensor Core, and the CUDA Programming Model#

4.4.2AMD GPU: CU, Matrix Core, and the ROCm/HIP Programming Model#

4.4.3Resource Balance Among General Control Flow, Non-Matrix Operators, and Matrix Throughput#

4.4.4SIMT, Cache/Shared Memory, and the Boundary of Programming Responsibility#

4.4.5Corpus Extensions: Biren, Hygon DCU, Muxi, and Moore Threads#

4.4.6Corpus Extensions: Tianshu Zhixin and Xiwang#