From GEMM to AI Accelerators
1.1.1Scalar Multiply-Add, Dot Product, Matrix Multiplication, and Tensor Contraction#
1.1.2GEMM, GEMV, Batched GEMM, and Operator Shapes#
1.1.3Computation Graphs for Forward Pass, Backward Pass, and Parameter Update#
1.1.4Compute Volume, Data Volume, Data Reuse, and Working Set#
1.1.5Latency, Throughput, Bandwidth, Capacity, and Energy#
1.1.6Execution Hierarchy from Processing Element to Chip#