3. CompilerCompilation and Program Mapping
Corpus alignment: use platform-specific compilation and loading contracts instead of assuming every accelerator exposes a CUDA-like stack.
3.1The AI Compilation Stack: Graphs, Operator Libraries, Kernels, and Device Execution63.2Frontend, Tracing, and Graph Capture53.3Intermediate Representation53.4Automatic Differentiation and Training-Graph Compilation53.5Graph-Level Optimization53.6Loop Transformation and Tensorization63.7Layout and Data-Movement Optimization53.8Memory Planning and Buffer Management53.9Kernel Code Generation and ISA Boundaries83.10Dynamic Shape and Runtime Specialization53.11Cost Models, Autotuning, and Auto-Scheduling53.12Distributed Compilation63.13The PyTorch Compilation Stack53.14JAX, XLA, and TPU Compilation73.15MLIR, TVM, and Deployment Compilation Stacks63.16NPU Compilation: Scratchpad, Engine Division of Labor, and Device Executables73.17Spatial Dataflow Compilation: Graph Mapping, Memory Placement, and Communication Scheduling73.18CIM/PNM Compilation and Memory-Centric Offload63.19Executable Artifacts, Loading Contracts, and Software-Stack Visibility63.20Compiler Validation, Debugging, and Case Studies6