AI Architecture & Systems

Decode Attention and KV Cache Kernels

2.10.1Paged KV Layout and Block Table#

2.10.2Reading and Reduction in PagedAttention#

2.10.3Split-KV and Long-Sequence Decode#

2.10.4Kernel Mapping of MHA/MQA/GQA#

2.10.5MLA Projection Absorption and Low-Rank Cache Access#

2.10.6KV Cache Append, Gather, Copy, and Compression#