AI Architecture & Systems

Matrix Computation Across Shapes and Scenarios

2.6.1GEMV and Small-Batch Decode#

2.6.2Small/Tall-Skinny GEMM#

2.6.3Batched GEMM and Grouped GEMM#

2.6.4Variable-Length Matrices, Dynamic Shapes, and Boundary Tiles#

2.6.5Low-Rank Matrices, LoRA, and Multi-Adapter GEMM#

2.6.6Weight Residency, Weight Packing, and Shape Specialization#