AI Architecture & Systems

4. ArchitectureAccelerator Architecture and AI Datacenters

Primary structure: Survey §2–§6. Core platform classifications follow the survey; additional cases are labeled as corpus extensions or contrasts.
4.1Scope, Terminology, and Evidence Standards for Architectural Comparison84.2AI Datacenter Context and Four Categories of Resource Pressure74.3Core Taxonomy: Four Classes of AI Accelerator Architecture74.4GPU: General-Purpose Parallel Execution and Specialized Matrix Computation64.5NPU: Heterogeneous Compute Engines and Shared Local Memory94.6Spatial Dataflow I: PE Arrays and Distributed Local Memory84.7Spatial Dataflow II: Reconfigurable Architectures64.8Spatial Dataflow III: Functional-Slice Streaming64.9Compute-in-Memory: SRAM, DRAM, and the Location of Computation84.10Compute Engines: Cross-Category Organization and Trade-offs74.11Memory Hierarchy: From Hardware Cache to Software Scratchpad74.12Programming Models: Understanding Hardware Constraints Through Software Layers84.13Packaging, Chiplets, and Logic–Memory Integration84.14Host, Remote Memory, and Memory-Centric Extensions74.15Physical Deployment Hierarchy and Scale-Up/Scale-Out Domains84.16Node-Scale Scale-Up Interconnect64.17Rack-Scale Scale-Up Interconnect64.18Pod-Scale Scale-Up Interconnect64.19Scale-Out, Network Datapaths, and Optical Interconnect74.20Topology Performance: Connectivity Is Not Effective Communication Performance74.21Collective Semantics: Endpoint Data Transformation74.22Collective Algorithms: Logical Communication Scheduling84.23Collective–Topology Mapping and Communication Offload74.24GPU Generational Evolution: Compute, Data Supply, and Cooperation Scope74.25Specialized Accelerator Generations: Workload Positioning and Hardware–Software Co-Design74.26Cross-Generation Comparison: Numerical Formats, Structured Sparsity, and Peak Conventions74.27Cross-Generation Comparison: Memory Hierarchy and Scale-Up Fabric84.28Power Delivery, Cooling, and Practically Operable Capacity94.29Future Design Challenges: Generality versus Specialization84.30Extensions and Contrast Cases: Photonics, Neuromorphic, and Limited-Disclosure Designs84.31Architectural Evaluation and Full-Stack Case Organization8