AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.3
From Naive GEMM to Tiled GEMM
2.3.1
Naive GEMM and Loop Reordering
#
2.3.2
CPU Cache Blocking and SIMD Microkernel
#
2.3.3
GPU Global-Memory GEMM
#
2.3.4
Shared-Memory Tiling and Register Blocking
#
2.3.5
Compute Reuse, Memory Coalescing, and Write-Back Optimization
#
2.3.6
Correctness Checking and Step-by-Step Performance Analysis
#
← Previous
2.2 Tensor Layout and Memory Access
→ Next
2.4 Tensor Core GEMM: Matrix Instructions, Data Layout, and Cooperation Scope