AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.10
Decode Attention and KV Cache Kernels
2.10.1
Paged KV Layout and Block Table
#
2.10.2
Reading and Reduction in PagedAttention
#
2.10.3
Split-KV and Long-Sequence Decode
#
2.10.4
Kernel Mapping of MHA/MQA/GQA
#
2.10.5
MLA Projection Absorption and Low-Rank Cache Access
#
2.10.6
KV Cache Append, Gather, Copy, and Compression
#
← Previous
2.9 Exact Attention Kernels
→ Next
2.11 Sparse and Irregular Kernels