AI Architecture & Systems

TPU v1 Case Study: From Systolic Array to GEMM Execution

SourcesTPU v1 paper, §2 and Fig. 1; later TPU generations are compared separately in Architecture.

1.7.1Matrix Multiply Unit and the Weight-Stationary Systolic Array#

1.7.2Unified Buffer, Weight FIFO, and Accumulator#

1.7.3Host Interface, External Weight Storage, and Instruction Control#

1.7.4Datapaths for Activation/Weight/Partial Sum#

1.7.5Coupling Matrix Multiply, Activation, and Output Write-Back#

1.7.6The Complete Execution Path of a Single GEMM Instruction#

1.7.7Boundary Between the TPU v1 Teaching Model and Later TPU Generations#