AI Architecture & Systems

Exact Attention Kernels

2.9.1Operator Decomposition and Memory Traffic of Standard Attention#

2.9.2FlashAttention: IO-Aware Tiling and Recomputation#

2.9.3FlashAttention-2/3 and Parallelism and Pipelining Optimizations#

2.9.4Causal Mask, Sliding Window, and Variable-Length Attention#

2.9.5Forward/Backward Attention Kernels#

2.9.6FlexAttention and Programmable Attention#