AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.6
Matrix Computation Across Shapes and Scenarios
2.6.1
GEMV and Small-Batch Decode
#
2.6.2
Small/Tall-Skinny GEMM
#
2.6.3
Batched GEMM and Grouped GEMM
#
2.6.4
Variable-Length Matrices, Dynamic Shapes, and Boundary Tiles
#
2.6.5
Low-Rank Matrices, LoRA, and Multi-Adapter GEMM
#
2.6.6
Weight Residency, Weight Packing, and Shape Specialization
#
← Previous
2.5 Advanced GEMM Pipelining and Scheduling
→ Next
2.7 Elementwise, Reduction, and Scan