AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.5
Advanced GEMM Pipelining and Scheduling
2.5.1
Double/Multi-Buffering and Software Pipelining
#
2.5.2
Asynchronous Copy and Producer–Consumer Synchronization
#
2.5.3
Warp Specialization and Role Assignment
#
2.5.4
Persistent Kernel and Persistent GEMM
#
2.5.5
Split-K, Stream-K, and Work Distribution
#
2.5.6
CTA Swizzle, Cluster, and Locality
#
2.5.7
Occupancy, Register Pressure, and Wave Quantization
#
← Previous
2.4 Tensor Core GEMM: Matrix Instructions, Data Layout, and Cooperation Scope
→ Next
2.6 Matrix Computation Across Shapes and Scenarios