2.4.1WMMA/MMA, Matrix Fragments, and Operand Layout#
2.4.2Instruction Tile, Warp Tile, CTA Tile, and Cluster Tile#
2.4.3ldmatrix, cp.async, and Operand Load Paths#
2.4.4Kernel Use and Synchronization of Hopper WGMMA/TMA#
2.4.5Blackwell tcgen05/TMEM and Paired-SM Kernels#
2.4.6Accumulator Layout, Numerical Precision, and Epilogue#
2.4.7Hardware Generation Selection, Fallback, and Operator Correctness#