AI Architecture & Systems

NPU Compilation: Scratchpad, Engine Division of Labor, and Device Executables

SourcesAWS Neuron and Huawei Ascend layer mappings.

3.16.1AWS neuronx-cc: Graph Compilation and Device-Specific Code Generation#

3.16.2The Distinct Entry Points of the NKI Compiler and the Graph Compiler#

3.16.3SBUF/PSUM Allocation, Prefetch, and Cross-Engine Scheduling#

3.16.4NEFF Artifacts, Runtime Loading, and Device Execution#

3.16.5Ascend: MindIR/ATC/Graph Engine and CANN#

3.16.6Operator Coverage, Static Shapes, Backend Constraints, and Model Porting#

3.16.7Separating Public Interfaces, Internal IR, and Inferred Information#