1. Microarchitecture
Survey alignment: §3.1–§3.2 and §5.1–§5.2. The GEMM, roofline, circuit, and RTL foundations remain the introductory curriculum.
1.1From GEMM to AI Accelerators61.2Numerical Representation, Arithmetic Circuits, and Precision Contracts81.3Execution Organization: SIMT, Heterogeneous Engines, and Spatial Execution81.4Processing Elements and Matrix Compute Units61.5Systolic Array Design71.6Dataflow and Data Reuse61.7TPU v1 Case Study: From Systolic Array to GEMM Execution71.8Registers and On-Chip Local State61.9On-Chip SRAM Management: Cache, Scratchpad, and Dedicated Buffers81.10DRAM and External Memory Interfaces71.11Data Movement and Asynchronous Execution71.12On-Chip Interconnect and Synchronization61.13Roofline and Performance Upper Bounds71.14Hierarchical GEMM Mapping and Tiling71.15Evolution of Matrix-Engine Datapaths and Control Granularity71.16Non-GEMM Operators and Dedicated Support61.17Hardware Modeling, Implementation, and Verification7